DeepSeek "R2" Officially Released — 90% Reduction in Inference Costs Reshapes the Landscape of Open-Source LLM Competition
機械翻訳 / Machine-translated
DeepSeek officially released its reasoning-specialized model "R2" on August 9, 2026 at 22:00 (UTC+8). The input token price of $0.014/M is approximately 89% cheaper than the current GPT-4o, while the company reports that R2 matches GPT-5 levels on major benchmarks (AIME 2025: 92.3%, SWE-bench Verified: 61.8%). The combination of open-weight distribution and commercial use availability represents a structural impact that could surpass the release of R1 in January 2025.
Model weights and a technical report were simultaneously published on GitHub and the official DeepSeek website. The architecture employs MoE (Mixture of Experts); the total parameter count has not been disclosed, but the active parameter count is described as approximately 40% lower than R1.
Inference costs (via API, as of August 9, 2026):
Benchmark preliminary figures (from the official DeepSeek report):
"R2's SWE-bench 61.8% is an internal test figure and hasn't been submitted to the official leaderboard yet. That said, running it locally feels noticeably faster than R1, and the code output is clean. I can't think of a better option considering the cost."
(X, engineering account, 2026-08-09)
In the 18 months since DeepSeek-R1 was released in January 2025, inference costs across the LLM market have dropped sharply industry-wide. As OpenAI, Anthropic, and Google have successively lowered their API prices, DeepSeek has continued its "asymmetric efficiency strategy." R1 surpassed 18 million downloads on Hugging Face within three months of release, and on-premise deployments by small and medium-sized businesses across Asia and Europe increased rapidly.
The license change is also significant. R1 had certain restrictions on commercial use, but R2 adopts the "DeepSeek Model License 2.0," which in principle allows free commercial use for services with fewer than 100 million monthly active users. This change is expected to accelerate the ripple effect on domestic SaaS vendors.
What an input price of $0.014/M tokens means in practice is that processing documents equivalent to one million characters (approximately 700K tokens) can be done for under $10. For large-scale document processing tasks in law, accounting, and healthcare, this represents a change that could shift monthly cost estimates by one to two orders of magnitude.
Because model weights are distributed directly rather than through an API, local deployment is immediately possible for industries such as finance, healthcare, and the public sector — sectors that "cannot send raw data to the cloud." In the Japanese market, where concerns about data sovereignty remain strong, the impact is expected to emerge earlier and more broadly than with other models.
AIME 2025's 92.3% is +12pt over R1, placing it in the same tier as models currently considered top-class. Developers who were considering integrating R2 into coding agents or mathematical analysis tools will immediately need to revise their cost projections.
The fact that R2 is an open-weight model from a Chinese company has not changed. Western companies that view security review and compliance costs as a barrier to adopting R2 will continue to exist. Japanese companies are now at a juncture where each organization must weigh "cost reduction vs. governance risk" on its own terms.
When R1 was released, competing companies revised their prices within two weeks. There is a strong possibility that some form of countermeasure will emerge within 72 hours this time as well.
What R2 is presenting is not a narrative of "the performance race is over and we've moved to a cost race." The current situation is one where both performance and cost are moving simultaneously. API providers that can no longer differentiate on either front are expected to lose their positions heading into the end of 2026.
The segment most significantly affected in the Japanese market is likely enterprise SaaS developers. If GPT-4o-level accuracy can be run on-premise at $0.014/M input, the economics of products that were "abandoned due to high AI costs" change entirely. Given the speed at which developers responded to R1, the evaluation cycle for R2 will likely run intensively over the next one to two weeks.
At the same time, the perspectives of data governance and export controls cannot be ignored. This is a moment where ROI should be calculated including the coordination costs with legal and information security departments; there will be cases where "let's just try it out" is not sufficient.
The two things to watch in the next 48 to 72 hours are: the timing of competing API price revisions, and confirmation of R2's official entry on the SWE-bench leaderboard. If the latter is confirmed, discussions around adoption in the coding agent market will move swiftly.
DeepSeek R2 has effectively nullified the binary trade-off between "performance or cost" in AI model selection. The combination of 90%-range accuracy and ▲89% inference cost could serve as a catalyst accelerating decision-making for companies considering AI adoption. The assumption within your organization that "we couldn't implement AI because of high costs" may have reached the point where it needs to be reconsidered, starting today.
※ This article was written by the AI writer (AI News) of the Mirai News editorial team.