Google "Gemini 2.5 Ultra" Officially Released — 2M Tokens × Scientific Reasoning Redefines the Foundations of Long-Document Analysis
機械翻訳 / Machine-translated
On August 9, 2026 (Pacific Time), Google DeepMind made its flagship model "Gemini 2.5 Ultra" generally available. With a context window expanded to an industry-leading 2 million tokens and a new SOTA score of 92.3% on GPQA (a scientific reasoning benchmark), the operational assumptions behind tasks requiring "long documents + high-precision reasoning" — such as legal document review, financial disclosures, and academic paper scrutiny — have begun to shift at a practical level.
Google DeepMind has begun general availability of Gemini 2.5 Ultra via Google AI Studio and Vertex AI. Compared to the previous-generation Gemini 2.0 Ultra, the context window has been doubled from 1 million to 2 million tokens. API pricing is set at $18 per 1M input tokens and $54 per 1M output tokens.
Shortly after launch, a wave of validation reports from practitioners appeared on X.
"I fed the entire 1,500-page case law database into Gemini 2.5 Ultra and it completed cross-referencing against specific clauses in 3 minutes. This changes work that used to take the legal team half a day."
The GPQA (Graduate-Level Google-Proof Q&A) score of 92.3% ranks as the highest among publicly available models. On the MATH benchmark (mathematical reasoning), it also recorded 97.1%, narrowly surpassing OpenAI o3 (96.4%).
Gemini 2.5 Ultra is positioned as the top-tier model in its series, following Gemini 2.5 Pro, which launched in March 2026. Whereas Pro was designed with inference speed and cost efficiency as priorities, Ultra is clearly differentiated as a flagship model that maximizes accuracy and context length.
A context window of 2 million tokens corresponds to approximately 2,000 pages of English text. This makes it possible to load large volumes of legal documents, financial reports, and research papers — which previously had to be "split and processed" — in a single pass, enabling cross-document comparison and contradiction detection.
Google simultaneously announced an upgrade to NotebookLM Pro powered by Gemini 2.5 Ultra, strengthening its integration with paper analysis workflows for researchers.
2 million tokens is not merely a spec competition. It enables treating an entire codebase exceeding 3,000 lines, or a full year's worth of contracts, as "a single context." The risk of "context breaks" caused by splitting content disappears, and the benefit is expected to be especially significant for documents with many cross-references.
A GPQA score of 92.3% is considered to represent a level at which "the majority of doctoral-level scientific questions can be answered correctly." The potential for use as an assistive tool in industries requiring specialized reasoning — such as drug discovery, materials science, and financial engineering — has expanded considerably.
This release also introduces fine-tuning options on Vertex AI. Adapting to industry-specific terminology and document styles becomes easier, and this can be seen as Google's full-scale entry into "platform competition for industry-specialized models," moving beyond the general-purpose model market.
Input at $18/1M tokens places this in the higher-cost tier even among current releases. However, the ability to process 2 million tokens in a single pass means there may be cases where the effective cost is actually lower compared to workflows that previously required multiple API calls. It is worth noting that fully utilizing 2 million tokens would result in an inference cost exceeding $36 per call.
Embedding within NotebookLM Pro has the potential to accelerate adoption among non-engineers. It serves as a pathway for legal, compliance, and research professionals to use long-document reasoning without going through an API.
What deserves the most attention with the release of Gemini 2.5 Ultra is not "the height of the scores" but rather "the doubling of context length." The task of "splitting content and restructuring prompts" — which has been a bottleneck in LLM workflow design — will become unnecessary for many use cases.
In industries such as legal, finance, and pharmaceuticals, where grasping the "full picture" of a document all at once is essential, this is not merely a convenience improvement — it means that the very premise of workflow design changes. Scenarios where designs that previously required multiple reasoning loops can be consolidated into a single inference call will clearly multiply.
On the other hand, cost must be designed carefully. A "just feed everything in" approach risks monthly costs far exceeding expectations. The core of any adoption decision lies in cost design: "which documents, at what granularity, and how frequently to process."
The availability of fine-tuning on Vertex AI is a move that draws the enterprise competition with Azure OpenAI firmly into a full-scale battle for industry-specialized models. We are now at the stage where the "accuracy × cost × context length" trade-offs relative to GPT-4o's $5/1M price point can be calculated concretely, industry by industry.
Gemini 2.5 Ultra carries significant practical impact not only along the axis of "further improvements in model performance," but also in terms of expanded design freedom through context length. The figure of 2 million tokens must be read not as "a data point on a spec sheet" but as "a variable in workflow design."
Two things will be the next focal points: whether OpenAI will directly compete on context length with the o4 series, and how far Vertex AI's fine-tuning pricing will support enterprise switching decisions. If real-world case studies of "long context × industry specialization" accumulate before year's end, the weight of adoption decisions could shift dramatically, all at once.
This article was written by an AI writer (AI News) from the Mirai News editorial team.