Google's "Gemini 2.5 Ultra" Officially Released — 2M Tokens × Video Understanding Reshapes Enterprise Analytics
機械翻訳 / Machine-translated
機械翻訳 / Machine-translated
On August 15, 2026, Google DeepMind officially released its flagship LLM, "Gemini 2.5 Ultra." The model's pillars are an industry-leading 2-million-token context window and native multimodal reasoning capable of processing up to four hours of video directly. It now competes within one to two percentage points of GPT-4.5 and Claude Opus 4, signaling a moment when the race among top-tier models has shifted from "benchmark gaps" to "depth of ecosystem integration."
The API became available on the same day via Google AI Studio and Vertex AI, with an enterprise tier for Google Cloud partners also released early. Key specifications are as follows.
On X (formerly Twitter), enterprise engineers in Japan have been sharing reactions like this:
"With Gemini 2.5 Ultra, you can feed an entire set of legal documents into 2 million tokens and run Q&A on them. This could fundamentally change the workflow of legal teams."
The Gemini series has been developed since Gemini 1.5 Pro in December 2024 with a focus on expanding context length and improving multimodal performance. In the first half of 2025, Gemini 2.0 Flash broadened enterprise adoption by emphasizing cost efficiency.
Entering 2026, OpenAI's GPT-4.5 and Anthropic's Claude Opus 4 launched in quick succession, intensifying competition at the top tier. Google chose to differentiate along two axes: "context length" and "video understanding."
The enterprise need that has become most apparent is "bulk processing of long-form content." Use cases that previously required chunked processing or RAG — such as contracts, case law, and video logs — can now be handled within a single prompt. This change directly benefits the bottom line by reducing pipeline design costs.
Two million tokens equates to roughly 1.5 million words in English, or approximately 6,000 A4 pages. Q&A can be completed while retaining an entire set of contracts, case law, and regulatory documents as context. A design in which "AI performs an initial review before a human reads it" becomes implementable at realistic costs in the legal and compliance domains.
Footage from cameras on manufacturing floors, in retail environments, and in medical settings can be fed directly into the API. Previously, a pipeline of video → text conversion (e.g., Whisper) → RAG was required, but these steps are now consolidated into a single API call. In terms of pipeline design with zero information loss, the foundation for on-site digital transformation has advanced by a level.
Across three metrics — MMLU, HumanEval, and MathBench — Gemini 2.5 Ultra sits in the 91–93% range, with the top three models clustered within two points of each other. Going forward, model selection criteria are expected to hinge less on benchmark differences and more on "integration with internal data pipelines" and "API cost structure."
The division of roles is now explicit: Gemini Nano on the device side (for Android/Pixel) and Gemini 2.5 Ultra on the cloud side. This announcement has further solidified design principles for a hybrid edge-cloud inference architecture.
At $1.50/Mtok for input, this is not cheap. Making full use of 2 million tokens drives up the cost per request. In production, "how much to pack into the context" becomes an unavoidable cost-design consideration.
The first thing worth noting is a structural shift: the race for context length has effectively reached a ceiling. Two million tokens is enough to cover the vast majority of enterprise use cases. The next competitive axis is expected to move toward latency, cost, and "flexibility in tool integration and fine-tuning."
Native video understanding integration is genuinely significant for on-site digital transformation in manufacturing, retail, and healthcare. Being able to perform inference directly on video data without converting it to text means eliminating information loss in the conversion process entirely. The benefits of this design change should surface in the numbers — in metrics like quality inspection accuracy and inventory management precision — one to two years from now.
On the other hand, with three models now in the same performance band, the decision-making process for "which one to choose" becomes more nuanced. The focus of model selection will likely move away from simple performance comparisons and toward TCO evaluations that encompass security policies, contract terms, and the cost of integrating with existing infrastructure.
Tokyo and Osaka region support effectively lowers the practical barrier for Japanese companies. The ability to handle compliance with Japan's Act on the Protection of Personal Information entirely within domestic regions is one of the conditions that will advance adoption consideration in financial, medical, and public-sector domains.
With the official release of Gemini 2.5 Ultra, the "context saturation point" of 2 million tokens has entered the toolkit of Japanese enterprise workflow design. For pipeline designers working with long-form documents, video, and multimodal content, this is an opportune moment to reconsider existing workflows. The next focal point is how OpenAI responds at this performance tier — the timing of a GPT-4.5 successor announcement and its pricing strategy are expected to be the fork in the road that determines the shape of the top-tier LLM competition for the rest of the year.
This article was written by an AI writer (AI News) from the Mirai News editorial team.