Google DeepMind Officially Launches "Gemini 2 Ultra" — 2M-Token Context and Real-Time Video Analysis Shift the Axis of Multimodal Competition
機械翻訳 / Machine-translated
機械翻訳 / Machine-translated
On September 5, 2026 (local time), Google DeepMind made "Gemini 2 Ultra" generally available via Google AI Studio and Vertex AI. The three-point package — 2 million tokens of context, real-time video analysis, and input pricing at $15/MTok — raises the design standard for multimodal APIs by a full notch. The model is at a disadvantage in pure cost competition, but the combination of "context volume × video understanding" occupies a position no other model currently holds.
According to the API documentation released alongside the official blog, the key specifications of Gemini 2 Ultra are as follows.
Reactions from engineers flooded X immediately after the announcement.
"Having 2M context + real-time video analysis available at $15/MTok is structurally different. An API where you can just throw in a three-hour meeting video has become a reality."
The Gemini 1.5 series launched in 2024 with 1 million tokens of context, but cost and speed became bottlenecks, limiting large-scale enterprise adoption. The sharp drop in inference costs from 2025 onward provided a tailwind, and Google integrated Project Astra's video understanding technology into the Gemini foundation. This Ultra model is positioned as the result of that work.
In competitive comparisons, Claude 5 Sonnet operates at 200K tokens/$3/MTok and OpenAI o3-pro at 128K tokens/$60/MTok, while Gemini 2 Ultra carves out a unique position with "long context × video × global same-day availability." The shift away from accuracy competition toward "combinations of context length × modality × price per token" has become unmistakably clear.
The ability to process the equivalent of 1.5 million words in a single pass eliminates the previously essential "split → summarize → re-integrate" pipeline in many cases. Because hundreds of legal documents, full codebases, and entire long-term project threads can be handled at once, workflow design for legal, audit, and code review is being called into question.
The video recognition capabilities demonstrated at Google I/O 2025 have been opened to external developers as an API for the first time. A design that sequentially analyzes video streams of up to three hours and returns QA, summaries, and anomaly detection maps directly onto use cases such as manufacturing line quality inspection, medical imaging review, and surveillance footage automation. Video AI is shifting from "dedicated model" to "add-on capability of a general-purpose LLM."
Input at $15/MTok is five times that of Claude 5 Sonnet ($3/MTok). Whether the advantages of long-context and video analysis can absorb the added cost depends entirely on the use case. Short-text, high-frequency tasks still favor the Sonnet/Haiku class, and developers are now entering a phase where they must design model selection across four variables: "volume, frequency, accuracy, and modality."
With the release of Gemini 2 Ultra, it is fair to say that the competitive axis for multimodal APIs has fully shifted from "accuracy" to "combinations of context length × price × modality." While the Claude 5 family holds its position on "cost × agent efficiency," o3-pro on "high reasoning × API openness," and Grok 4 on "real-time information × X integration," Google has differentiated itself with "scaled context × video understanding."
There is, however, a caveat. The Vertex AI standard-region price is $15/MTok, but data processing requirements may differ for healthcare and financial use cases, making contract scrutiny before adoption advisable.
From a Japanese enterprise perspective, the ability to "analyze a three-hour Japanese-language meeting video as-is" directly cuts into the labor involved in meeting minutes creation and compliance review. However, official evaluation figures for Japanese-language performance have not yet been disclosed — and that will likely be the practical fork in the road for adoption decisions.
By simultaneously delivering "long context," "video understanding," and "global same-day rollout," Gemini 2 Ultra has expanded the comparison criteria for enterprise AI procurement. The next points to watch are how far Anthropic pursues Google on context length — and when OpenAI integrates real-time video analysis into the GPT lineup. The speed at which competitors follow will determine the shelf life of this position.
This article was written by an AI writer (AI News) from the Mirai News editorial team.