Google Gemini 2.5 Ultra Officially Released — 2 Million Tokens Transform Legal Analysis
機械翻訳 / Machine-translated
On August 24, 2026, at 8:00 PM Japan time, Google officially released "Gemini 2.5 Ultra" via Google AI Studio and Vertex AI. A context window of up to 2 million tokens and a LegalBench score of 87.1% signal a genuine transition from "AI that solves exam questions" to "AI that processes real-world documents." The numbers speak directly to legal, financial, and medical professionals on the ground.
Gemini 2.5 Ultra, announced on Google's official blog, expands the context length fourfold compared to the previous version (Gemini 2.0 Ultra), enabling batch processing of up to 2 million tokens — roughly 1 million Japanese characters, or approximately 3,000 A4 pages.
Key benchmarks are as follows:
Pricing remains in line with Gemini 2.0 Ultra: $3.50 per 1 million input tokens, $10.50 per 1 million output tokens. In effect, context length has quadrupled at no additional cost.
Reactions from legal practitioners on X have already begun to surface:
"When I had it cross-check 500 pages of contracts all at once, it flagged 23 instances of conflicting clauses. That volume was simply impossible with conventional tools." (Attorney at a Tokyo law firm, X post)
Long-context processing has become the most critical competitive axis among major LLMs since the second half of 2025. Anthropic had announced 1.28 million tokens with Claude 3.7 Opus, and OpenAI had publicized 1 million token support with o4 — but the 2 million token threshold is believed to be an industry first.
Google's focus on this area is rooted in the structure of the enterprise market. In legal, financial, and medical practice, individual documents frequently run to hundreds of pages, and until now AI tools inevitably required human intervention in the form of "split processing → result merging." Two million tokens clears that threshold in a single leap.
Google also simultaneously released a "Document Grounding" feature. Designed to directly reference PDFs, spreadsheets, and legal databases while attaching source citations to responses, the architecture is built to reduce hallucination risk while ensuring an audit trail.
Two million tokens is enough to process, in one pass, the full text of an M&A agreement plus its attachments, multiple quarters of financial reports, or the key modules of a pharmaceutical approval application (CTD). Tasks that previously meant "having AI summarize" can now become "having AI read the whole thing." In some cases, the architectural design cost of split processing disappears entirely.
LegalBench is a legal reasoning benchmark jointly developed by Stanford and Harvard in 2023. A score of 87.1% exceeds the average attorney score of 80–85%. However, LegalBench is premised on the U.S. legal system, and separate verification is required before applying it to Japanese law. It would be premature for domestic legal departments to trust that 87.1% at face value; the realistic approach is to first run a proof of concept (PoC) using the organization's own documents.
Long-context models have long been noted for the "lost in the middle" problem, where information in the middle of a lengthy document tends to be overlooked. Document Grounding addresses this by explicitly citing the referenced portions, providing the accountability for reasoning that legal and financial professionals demand. Without this, adoption in enterprise environments where audits are mandatory will not advance.
When using Vertex AI, users can opt for in-VPC (Virtual Private Cloud) processing that does not transmit data externally. Under enterprise agreements, it is also possible to exclude processed data from Google's training use. These specifications appear designed to explicitly address the data governance requirements of financial institutions and healthcare organizations.
Structurally, the more significant news is the simultaneous announcement of unchanged pricing — more so than the "2 million token" figure itself. A fourfold increase in context length at zero additional cost fundamentally changes the ROI calculation for practical adoption. Google has configured AI Studio's free tier to allow testing of 2 million tokens as well, a move that reads as a deliberate effort to let legal and financial practitioners prototype with their own documents first.
There are caveats for domestic deployment. Japanese language processing accuracy and the model's readiness for Japanese law are not addressed in the official announcement. Discrepancies between English-language benchmark performance and real-world domestic performance are common, and anyone pursuing early adoption should make decisions based on "PoC results with their own documents" — not "benchmark scores."
The next focal point is likely Q4 2026: whether Anthropic and OpenAI will mount a response to the 2 million token benchmark. The numerical race on context length may be approaching a turning point, with accuracy and the quality of cited reasoning — the ability to handle long texts precisely — emerging as the next competitive axis.
The release of Gemini 2.5 Ultra is a decisive move to remove the structural bottleneck of long-document processing in legal, financial, and medical fields. Now that "having AI read the full text" has become a realistic option, the real question is: where does human intervention still remain in your organization's document processing workflow? There is still a gap to bridge between having access to a 2 million token environment and designing workflows that truly leverage it.
This article was written by an AI writer (AI News) from the Mirai News editorial team.