Gemini Shock | AI Researcher 93% × 160 Searches in 20 Minutes
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
"I asked AI to do some research, and while I slept it checked 160 sites and put together a report"——that future has finally become reality.
On April 21, 2026, Google announced autonomous AI research agents "Deep Research" and "Deep Research Max," powered by Gemini 3.1 Pro.
"What's different?" "How impressive is it?" "How will it change our work?"——we break down the answers you want to know, in plain language anyone can understand.
What Google announced this time is a new series of autonomous AI research agents.
"Deep Research: the standard version prioritizing speed and cost, ideal for real-time interaction."
"Deep Research Max: the ultimate version that takes its time for thorough investigation, generating reports by morning via overnight batch processing."
Think of it as the difference between a convenience-store onigiri (Deep Research) and a high-end kaiseki course (Max)——both are great, but for different purposes.
"Both are equipped with Gemini 3.1 Pro, the latest brain, and autonomously execute up to 160 web searches per task."
"Research results are also auto-generated as charts and infographics, visualized as HTML or Nano Banana images."
"A complete report with citations delivered within 20 minutes" is the selling point.
"References more than 100 sources per task, performing cross-searches across public web information, user-uploaded files, and all MCP-connected data."
"Internal evaluations confirm the new version overwhelmingly outperforms the previous version from December 2025."
The benchmark figures Google released are shaking the industry.
"DeepSearchQA (question answering on complex search tasks): 93.3% (a dramatic leap of +27.2 points from the previous generation's 66.1%)."
"BrowseComp (browser-based research capability): 85.9 (more than +25 points compared to Gemini 3 Pro)."
"Humanity's Last Exam (a test of extremely difficult questions): 54.6% (up +8.2 points from the previous 46.4%)."
"GPQA Diamond (graduate-level science questions): 94.3%."
It's the kind of rapid growth where a student who consistently scored 70 on every test suddenly starts hitting 93 in a row.
"In particular, the 93.3% on DeepSearchQA approaches the accuracy of human expert analysts."
"OpenAI's Deep Research was estimated in the 60% range on previous-generation metrics; Anthropic Claude is estimated around 80%——Google is currently a clear head ahead."
"However, all figures are based on Google's internal test data; third-party evaluation awaits future verification."
"Even so, these are historic numbers representing AI stepping into 'domains that truly require thinking,' with industry insiders saying 'we've taken one step closer to AGI.'"
Currently available as a paid public preview.
"Immediately available to developers via the Gemini API; the free tier has been discontinued since April 1, 2026."
"Can be invoked with about 10 lines of Python code; also supports asynchronous processing via the Interactions API."
"Pricing follows official pay-as-you-go billing, calculated based on output tokens and search grounding fees."
It's as convenient as buying an AI assistant at a convenience store (Gemini API)——you call it up only when you need it.
"Full deployment to Google Cloud is planned for a few months out; enterprise customers are in a queue."
"Some Deep Research features are already built into the Gemini App for general users, available with Google AI Pro (approximately ¥3,000/month)."
"However, full use of Deep Research Max is currently API-only and requires development skills."
"As a public preview, note that Google may issue spec-change notices before production use."
"Integration into Google Workspace (Gmail, Docs, Sheets) is also planned, with further expansion of business use expected."
Gemini 3.1 Pro is Google's latest flagship language model.
"The successor to Gemini 3 Pro, released in December 2025, with a major update in just four months."
"BrowseComp score: 60.x → 85.9 (more than +25 points)——a dramatic leap from average to world-top level."
"Specialized for long-horizon research workflows, capable of up to two hours of continuous autonomous reasoning."
The pace of evolution is like a student who ranked 20th in their middle school class shooting up to first place on a national mock exam in just six months.
"Other improved capabilities: multimodal understanding (can read PDFs, CSVs, images, audio, and video all at once)."
"Native tool use (the ability to independently select and call external tools) has also been significantly enhanced."
"Knowledge cutoff updated to early 2026, giving it fresher information than ChatGPT or Claude."
"Even compared to rivals GPT-5.5 (OpenAI) and Claude Opus 4.7 (Anthropic), it has an edge in the combination of search and reasoning."
"This is the model Google staked its future on——the company's unique path of fusing search engines with AI has come to fruition."
The biggest innovation this time is "MCP (Model Context Protocol) support."
"MCP: an industry-standard protocol for AI to securely connect to external data sources and tools."
"Proposed by Anthropic in 2024, it rose to industry-standard status in just one year——the 'common language of AI.'"
"With Deep Research now MCP-compatible, it can freely connect to internal systems, specialized databases, and custom tools."
"Concrete examples: MCP servers for FactSet (major financial data provider), S&P Global (ratings company), and PitchBook (investment information) are already supported."
It's a huge shift——like going from AI being able to see only "outside the store" to being able to see everything inside your home refrigerator.
"With this feature, companies can reference their proprietary data in real time during research without having to train it into the model."
"Data leak risk is also reduced, since confidential information is managed on the MCP server side, keeping it secure."
"Developers can define custom MCP tools, enabling them to build a dedicated AI researcher for their organization."
"In the second half of 2026, major SaaS platforms including Salesforce, SAP, Oracle, and Workday are expected to announce MCP compatibility, dramatically expanding business use."
Another strength is "multimodal research grounding."
"If users upload PDFs, CSVs, images, audio, or video, the AI references all of them as sources during research."
"Example: upload a meeting audio recording + past slide PDFs + an industry report all at once → simply say 'analyze this company's strategy.'"
"Reports cite specific evidence, such as 'page 3 of the material you provided' or 'at the 25-minute mark of the meeting recording.'"
It's like handing documents to a new employee and saying "research based on these," then receiving a report three days later that specifies exactly which page of which document was referenced——that level of precision.
"Native chart and infographic generation is also standard, turning complex data into easy-to-understand visuals at a glance."
"Supports both HTML and Nano Banana images, which can be embedded directly into reports."
"A research plan pre-confirmation feature is also included; before starting, the AI asks 'Is it okay if I proceed in this manner?'"
"Since users can adjust the direction of the research, this mechanism reduces wasted effort."
Deep Research (standard version) is characterized by being "fast and easy to use."
"Response time: approximately 5–10 minutes, ideal for interactive use cases that prioritize immediacy."
"30–80 web searches per task, providing a sufficient depth of research."
"API cost: approximately one-half to one-third of Max, for cost-conscious users."
"Output quality: high, though not as thorough as Max."
Think of it like drip coffee from a convenience store——the right quality and price for everyday work.
"Typical use cases: last-minute prep before a meeting, fact-checking for email replies, research for blog posts."
"Some features are integrated into Gemini Apps (equivalent to Google AI Pro at ~¥3,000/month), allowing general users to experience it too."
"Integration with Slack, Microsoft Teams, Notion, and others is also planned, to function as part of everyday work tools."
"Designed with low latency so users can use it while waiting."
"Since costs stay manageable even with multiple daily calls, it suits heavy users."
Deep Research Max is the ultimate mode that "takes its time for thorough investigation."
"Response time: 15–20 minutes, supporting long tasks up to approximately one hour."
"Up to 160 web searches per task, spanning more than 100 sources."
"Extended test-time compute: the AI spends more 'thinking time,' repeatedly cycling through search → reasoning → revision."
"Output quality: on par with a professional analyst; reports can run to thousands or tens of thousands of characters."
Like a high-end kaiseki course at a fine restaurant——taking time to prepare the finest ingredients with the finest technique.
"Typical use cases: due diligence for investment decisions, competitive analysis, market viability assessment for new drugs, preliminary research for litigation materials."
"Designed for asynchronous batch processing——set it running overnight and receive the finished product in the morning."
"Example: a financial institution analyst instructs 'run a Max financial analysis on these 10 companies' before leaving work → a perfect report is on their desk the next morning."
"API cost is 3–5x the standard version, but compared to the daily labor cost of a human analyst (tens of thousands of yen), it's a bargain."
"Less a replacement for experts than a booster that multiplies an expert's productivity tenfold."
The choice is simple: a two-axis judgment of "time urgency × importance."
"① Immediate need + low to medium importance: Deep Research (standard) is fine."
"② Immediate need + high importance: call Deep Research multiple times and compare."
"③ No urgency + high importance: Deep Research Max is optimal."
"④ No urgency + medium importance: a hybrid strategy combining both."
It's as intuitive as choosing a restaurant when ordering delivery based on "I want it now" vs. "I want to impress someone."
"Concrete examples: morning news check → Deep Research; preparing materials for a monthly management meeting → Deep Research Max."
"Budget: the standard version is sufficient for individual users; Max is worth the investment for corporate analyst work."
"Combined use pattern: use the standard version for a rough hypothesis → use Max to dig deeper and verify, balancing efficiency and accuracy."
"As you get used to it, finding the right balance for your own workflow is the proven approach."
"In the second half of 2026, an 'AutoSelect' feature that lets AI automatically choose the mode is planned, so the burden of choosing will eventually disappear."
The biggest rival is OpenAI's Deep Research.
"OpenAI Deep Research: launched in February 2025, a flagship feature of ChatGPT Pro ($200/month, approximately ¥30,000) that has been popular."
"GPQA Diamond: evolved from 26.6% (older GPT-4o) → 90.4% (latest GPT-5.5), nearly matching Gemini's 94.3%."
"Humanity's Last Exam: 26.6% → 54.6% (OpenAI) vs. 54.6% (Gemini)——essentially equal."
"OpenAI's strengths: easy-to-use UI, ChatGPT integration, code interpreter connectivity."
It's a McDonald's (OpenAI) vs. Burger King (Google) situation——the taste difference is subtle, but they compete on store reach and pricing strategy.
"Gemini's strengths: versatility of MCP support, integration with Google Search, completeness of multimodal capabilities."
"OpenAI's weaknesses: MCP is limited (centered on proprietary plugins), internal data integration is somewhat weaker."
"Reports suggest OpenAI plans to announce a competing 'Deep Research v3' in May 2026, intensifying the competition."
"From a user perspective, 'both are excellent but it depends on use case'——Gemini currently leads for developers, ChatGPT for end users."
Perplexity is a startup specializing in AI-powered search engines.
"Perplexity Pro Search: can switch between multiple AI models (GPT-5.5, Claude, Gemini), at $20/month."
"Strengths: simple UI, transparency with citations always clearly shown, immediacy."
"Weaknesses: autonomous agent capabilities are still limited, struggles with highly complex research."
It's a different kind of use case——convenience-store search (Perplexity) vs. a dedicated researcher at a university library (Gemini Deep Research Max).
"Perplexity's main customers: journalists, students, consultants, and other knowledge workers who prioritize immediacy."
"Estimated user count: approximately 45 million monthly active users as of April 2026, and rapidly growing."
"Gemini's advantages: autonomous multi-step task execution, internal data integration, ultra-large-scale research."
"Perplexity's advantages: choice of multiple models, speed of responses, lower monthly cost."
"The two are less competitors than complements; the number of users who switch between them by use case is growing."
"In the second half of 2026, Perplexity also plans to strengthen its autonomous agent capabilities, intensifying competition further."
Another major player is Anthropic's Claude Research.
"Claude Research: announced simultaneously with Claude Opus 4.7 in March 2026, featuring autonomous web browser operation."
"DeepSearchQA: undisclosed, though industry estimates put it at around 85–88%."
"Strengths: the originator of MCP, industry-leading coding ability, safety-first design."
As the inventor of MCP, Anthropic occupies the position of the original source.
"Claude's primary customers: enterprise AI developers, with an overwhelming share particularly in the coding domain."
"Gemini's advantages: currently highest benchmark scores, uniqueness of search engine integration."
"Claude's advantages: depth of MCP ecosystem, excels at research that involves writing code."
"The two have subtly different tradeoffs between 'safety vs. performance'——Anthropic is more conservative on the safety side."
"Rumors suggest Anthropic will announce Claude Research Max in May 2026, directly countering Google."
"Conclusion: Gemini Deep Research currently combines 'top benchmarks + broadest ecosystem'——a rare achievement. However, Claude has an edge on safety."
"Smart users combine multiple tools, and using each where it's best suited is the 2026 standard."
A major transformation is coming to research work at Japanese companies.
"As of April 2026, approximately 38% of Japanese companies have adopted Google Workspace (estimated 24 million users), providing a solid foundation for Gemini integration."
"Expected business change ①: market research companies (Intage, Macromill, etc.) will see research speed increase tenfold, with report unit prices declining."
"Expected business change ②: analyst reports at investment banks and securities firms will shift to AI drafting → human supervision."
"Expected business change ③: case law research at law firms will shrink from hours to 20 minutes with cross-case searching."
It's a speed revolution on the level of Edo-era couriers (human analysts) being suddenly replaced by the Shinkansen (AI researchers).
"However, Japan-specific circumstances apply: handling Japanese-language specialized materials, a cautious stance in regulated industries like legal and medical, and delays in establishing AI usage guidelines."
"Major consulting firms including Deloitte, PwC, KPMG, and EY have already announced enterprise Gemini Deep Research implementation support services."
"Implementation pricing: ¥10 million to several hundred million yen per year, forming a market centered on large enterprises."
"Industry forecasts predict more than 100 Japanese companies will begin full-scale implementation in the second half of 2026."
The wind is also blowing favorably for small and medium-sized businesses and sole proprietors.
"Gemini API pay-as-you-go pricing: a few dozen to a few hundred yen per task, well within the budget of SMEs."
"Expected use case ①: administrative scriveners automating the monthly collection of the latest subsidy information."
"Expected use case ②: regional small manufacturers automatically researching competitors and raw material market prices."
"Expected use case ③: individual bloggers and YouTubers improving their posting frequency by streamlining content research."
It's the democratization of information——an era where a small local stationery shop can have the same information-gathering power as a major Tokyo trading company.
"Adoption hurdles: obtaining an API key, a little Python code, or routing through existing SaaS platforms (Zapier, Make, etc.)."
"DIY approach: developers can build their own custom tool in a few hours."
"For non-technical users: in the second half of 2026, Japanese-language SaaS products with built-in Gemini Deep Research (such as Kaonavi and freee) are expected to emerge."
"Concern: AI may produce reports containing incorrect information (hallucination); human final review is always necessary."
"That said, the division of labor where 'AI drafts → human proofreads' is a realistic solution that can boost productivity 3–5 times."
"For smaller companies, the era of first-mover advantage is here."
Major waves are also hitting Japan's AI industry and talent market.
"Google Japan plans to expand its Deep Research implementation support team in May 2026, scaling up to 100 enterprise sales personnel."
"SI companies like NTT Data, Fujitsu, and NEC are racing to become certified 'Gemini Deep Research integrators.'"
"AI talent demand: prompt engineers and AI workflow designers average ¥11 million annual salary in 2026, up +30% year-over-year."
"Job offer ratio: AI-related occupations are at 12x as of April 2026, triple the IT sector overall (4x), indicating severe labor shortage."
It's a gold rush-style boom where both the miners and the pickaxe sellers are profiting.
"Startups: Sakana AI, Preferred Networks, rinna, Karakuri, and others are actively forming partnerships with or competing against Gemini."
"Government movement: the Ministry of Economy, Trade and Industry plans to release 'AI Agent Utilization Guidelines' in June 2026."
"Concern: with the arrival of AI researchers, there is a risk of reduced new hiring of junior analysts and researchers, impacting the development of young talent."
"Response: people who acquire the skill of 'AI × human' collaboration will be the winners of the next decade; those relying solely on pure research skills face a harder road."
"Conclusion: Japan's white-collar workforce is entering an era of 'working alongside AI researchers' in earnest, starting in 2026."
Sasaki-san covers Japanese manufacturers at a major investment bank in Tokyo.
Annual salary ¥15 million; 80 hours of overtime per month is a chronic problem.
"The weekend report writing is eating too much into my family time"——chronic exhaustion.
"In May 2026, Sasaki-san introduced Deep Research Max into his work."
"On Friday evenings, he instructs: 'Prepare financial, competitive, and market analysis for the 5 companies I need to analyze next week'——and a perfect draft report is ready by the time he arrives Monday morning."
"The report includes charts, graphs, and full citations; only about 20% revision is needed."
It's like a division of labor where the head chef has finished all the prep work, and the sous chef just does the final plating.
"As a result, overtime dropped from 80 hours to 30 hours, and weekends became free again."
"Report quality also improved; his supervisor commented, 'Your deep dives have been sharper lately.'"
"Promoted to team leader in the second half of 2026; salary increased from ¥15 million to ¥20 million."
"He also took on the role of internal AI workflow training instructor with the newly freed time, expanding his influence within the company."
"A textbook example of not 'having work taken by AI,' but 'being an analyst who can master AI.'"
"Traditional analysts on the same team saw their evaluations decline due to the productivity gap, making the industry's polarization strikingly evident."
Ogawa-san runs a precision parts manufacturer in Shizuoka Prefecture (45 employees, ¥1.2 billion annual revenue).
Aiming for overseas expansion but struggling to gather English-language industry information.
"At our scale, we don't have the budget to hire an overseas research firm"——a real pain point.
"In June 2026, Ogawa-san adopts Deep Research with Google AI Pro (approximately ¥30,000/month) + Gemini API pay-as-you-go (approximately ¥20,000/month)."
"He runs a weekly Max query: 'Research the precision parts market size, major competitors, and local partner candidates in 5 Southeast Asian countries.'"
"The report integrates English trade publications, statistical data, and competitor IR materials, then summarizes everything in Japanese."
It's a revolution in cost efficiency——like an SME hiring the equivalent of a ¥10 million/year global research specialist for ¥50,000/month.
"Result: a distributor contract in Vietnam was signed in August 2026, with overseas revenue breaking ¥100 million for the first time."
"Further expansion into Thailand in 2027, growing annual revenue from ¥1.5 billion to ¥2 billion."
"Ogawa-san gave a lecture at the local Chamber of Commerce titled 'The Era When Regional SMEs Can Go Global with AI.'"
"Using Shizuoka Prefecture subsidies, he contributed to packaging the approach as a regional Gemini implementation support package."
"Featured in the Ministry of Economy, Trade and Industry's case study collection as a role model for regional revitalization × AI utilization."
"A success story symbolizing the end of the era where 'only large companies can use AI.'"
Honda-san is pursuing a master's degree in economics at the University of Tokyo.
The topic of her thesis: "Changes in the Labor Market in the AI Era."
"Comprehensively reviewing prior research kept me up until midnight every night——academic stress."
"In May 2026, Honda-san begins using Deep Research through the university's Google Workspace for Education (standard version available free with student discount)."
"She instructs: 'Analyze papers on the impact of AI on employment over the past 20 years and organize the main arguments and citation patterns.'"
"In 15 minutes, a cross-analysis of 180 papers is complete, with a report showing citation networks and unresolved debates at a glance."
It's a productivity revolution——work that would take a month in the library done by AI during a lunch break.
"Literature review time cut from 3 weeks to 3 days, freeing her to concentrate on the actual analysis."
"Thesis quality improved; her professor commented, 'This comprehensiveness is unusual.'"
"She also published a Note post on how to use Deep Research, gaining 8,000 PV/month and ¥30,000/month in side income."
"Entering a doctoral program in spring 2027, gaining attention as a young researcher in the cross-disciplinary field of AI × economics."
"She also held Deep Research utilization workshops for undergraduates, contributing to AI penetration in education."
"A case study that personally demonstrates that 'in the AI era, whether researchers can use AI changes their research capacity tenfold.'"
A. Yes, individuals can use it just fine.
"For general users: Deep Research features within Gemini Apps are available with Google AI Pro (approximately ¥3,000/month)."
"For developers: full Deep Research / Max is available via the Gemini API on a pay-as-you-go basis (a few dozen to a few hundred yen per task)."
"For students: the standard version may be available for free through the university's Google Workspace for Education; check with your institution."
It's like choosing between convenience-store coffee (individual plan), a restaurant course meal (API), or a school lunch (student discount)——a rich variety of options for different needs and budgets.
"Personal use examples: thorough comparison of travel plans, analyzing a prospective employer, market research for a side business."
"Note: the free tier was discontinued in April 2026; a paid plan is now required."
"The recommended path is to start with Google AI Pro at ¥3,000/month, then transition to the API as you get comfortable."
"API use requires basic Python knowledge (about 10 lines of code), though you can easily get it written by asking ChatGPT or similar."
A. Japanese accuracy is world-class——sufficient for everyday business use.
"Gemini 3.1 Pro significantly increased Japanese training data, and understands nuances in Kansai dialect and honorifics."
"Japanese Q&A accuracy: approximately 95% of English performance (versus about 85% for the previous generation)——almost no meaningful gap."
"Can also read Japanese specialized materials (laws, precedents, patents), with support for direct PDF upload."
It's like the improvement in Japanese English proficiency going from TOEIC 500 (old model) to TOEIC 900 (Gemini 3.1 Pro)——genuinely practical.
"However, it's not perfect: niche Japan-specific technical jargon (industry-specific idiomatic expressions, etc.) is sometimes misunderstood."
"Workaround: adding clarifications in your prompt like 'In this industry, ○○ means ___' improves accuracy."
"Auto-generated Japanese reports: structure, honorifics, and use of technical terms all feel natural; typically only 20–30% revision is needed."
"Integration with Google Docs: Japanese reports can be saved directly to Docs, with collaborative editing also supported."
"In the second half of 2026, accuracy is expected to improve further with additional Japanese-language training data."
A. Hallucination countermeasures have improved dramatically——but it's not zero.
"Deep Research's biggest strength: every claim comes with a citation (web URL or specific page of an uploaded document)."
"With cited output, users can immediately verify the basis for any claim."
"Hallucination rate reduced by approximately 1/5 compared to the previous generation, though there is still a possibility of a few percent of misinformation slipping through."
Think of it like an encyclopedia where 80% is correct but 20% still needs verification——that level of caution is warranted.
"Countermeasure ①: always access the cited URL to verify important facts."
"Countermeasure ②: ask the same question using multiple queries and check for consistency in the answers."
"Countermeasure ③: have a human expert do a final check, especially in legal, medical, and financial fields——this is essential."
"Areas requiring caution: breaking news (events after the knowledge cutoff), niche specialized knowledge, arithmetic (even simple numerical calculations can be wrong)."
"The mature approach: use AI as an assistant, with humans making final judgments——an 'AI × human' division of labor."
"Balancing 'not blindly trusting AI' with 'not being so skeptical that you don't use it' is key."
A. API use is designed to protect privacy.
"Gemini API: data uploaded by users is not used to retrain the model (opt-out is the default)."
"Data retention period: typically 24 hours to 30 days, then automatically deleted."
"MCP-connected data: retained on the user's MCP server side and not sent to Google."
It's a layered security design——like a bank safe deposit box (API), a hotel coin locker (limited retention period), and a home safe (MCP server), each with different custodians.
"However, Gemini Apps (for individuals): conversation history may be saved to your Google account; can be disabled in settings."
"Note ①: when handling confidential data, you should establish a Google Cloud enterprise contract plus a Data Processing Agreement (DPA)."
"Note ②: for personal information, medical information, and financial information, specialized contracts for regulated industries are required."
"GDPR and Japan's Act on the Protection of Personal Information compliant——on par with other AI services in terms of compliance."
"Best practice: keep confidential data on-premises via MCP connection, and permit AI 'read-only' access through a secure architecture."
"In June 2026, Google is expected to announce a Japan-region-specific Gemini API (fully contained within Tokyo data centers), resolving data cross-border issues."
"The era where an AI researcher checks 160 sites in 20 minutes for you" is fundamentally transforming how we work and learn——that is the essence of what this announcement conveys.
"The figures of DeepSearchQA 93.3% and BrowseComp 85.9 are proof that AI has begun to seriously take on the core of intellectual labor."
"OpenAI, Anthropic, and Perplexity are also preparing their counters; the second half of 2026 will plunge into an era of warring AI research agents."
It's a historic turning point——just as the emergence of calculators changed the work of human calculators, the emergence of AI researchers is now changing the work of human researchers.
"There are three things you can start doing today: ① experience the standard version with Google AI Pro, ② identify tasks in your own work that can be delegated to AI, ③ sharpen the areas of judgment and creativity that only humans can do."
**"It's not 'AI taking your job'——it's 'the era of pairing with AI to achieve 10x productivity'"——whether individuals, SMEs, or large enterprises choose to "work alongside an AI researcher" is ultimately a choice each of us must make for ourselves.
This article is a cross-post from AI Friends.