DeepSeek V4 Officially Released — 671B MoE × 90% Inference Cost Reduction Shifts the Competitive Landscape of Open-Source LLMs
機械翻訳 / Machine-translated
DeepSeek officially released "V4" on August 15, 2026. The model adopts a Mixture-of-Experts (MoE) design with 671B total parameters while limiting active parameters during inference to 67B, achieving scores on par with GPT-4o on MMLU and HumanEval while cutting API costs by approximately 70% compared to V3. This marks what is effectively the first time a MoE architecture has been released under an MIT license, and the axis of open-source LLM competition has shifted once again.
Model weights and an API were simultaneously released on the official GitHub repository and Hugging Face. Key specifications are as follows:
Benchmark scores (as reported by DeepSeek) are MMLU 91.3%, HumanEval 88.7%, and MATH 74.2%, claimed to be on par with or surpassing GPT-4o. Independent third-party verification has not yet been fully compiled.
"I ran DeepSeek V4 locally and it feels almost identical to GPT-4o in practice. The API cost is dramatically cheaper. This will change enterprise procurement decisions." (AI engineer, 42,000 followers)
DeepSeek also shocked the market with its low inference costs at the release of V3 at the end of 2024, prompting companies in Japan and Europe to begin reconsidering their procurement choices. V4 further improves upon V3's MoE efficiency, achieving both speed and cost reduction through a design that "limits active parameters used during actual inference to 67B."
As of 2026, with global electricity consumption reduction a pressing concern, demand for MoE — which delivers "the same performance with less compute" — is also growing from the infrastructure side. DeepSeek V4 can be said to directly address that trend.
The input unit price of $0.014/1M is estimated to be roughly one-seventh that of GPT-4o. For an application processing 10 million tokens per month, the monthly cost works out to approximately $140 — a figure that will realistically influence adoption decisions by SaaS-oriented startups.
MoE has been adopted by Gemini 1.5 and Mixtral, but "training instability" and "inference infrastructure complexity" had been barriers to self-hosting. With V4 releasing its weights, the number of on-premises deployment cases is expected to surge rapidly.
Regarding the adoption of models developed in China, a trend has continued in Japan since 2025 — particularly among finance- and defense-related companies — of codifying "prohibition of business use of Chinese-made LLMs" into internal regulations. Decisions to adopt V4 are expected to vary greatly depending on industry and data sensitivity.
What makes V4 significant is not simply that "another cheap model has arrived," but rather the structural shift that MoE architecture has now reached a practical level as open source. Until now, full-scale MoE deployment had been limited to internal implementations by major cloud providers. The release of weights creates, for the first time, a situation in which companies and research institutions that prioritize data sovereignty can operate this architecture under their own control.
In the Japanese market, many companies — given strong privacy concerns — hold data they "do not want to send to the cloud." More than the scale of cost reduction (up to 90%), we believe the fact that "an option exists to run a high-performance model under complete in-house control" will become the starting point for procurement discussions.
However, the benchmark figures published by DeepSeek are still awaiting independent verification. Third-party evaluations are expected to emerge within 72 hours of release, and the results there will likely cause assessments to swing significantly.
DeepSeek V4 is not merely an update in a cost competition — it represents a structural change that brings MoE architecture to practical use as open source. The next focal points are (1) independent performance verification through third-party benchmarks, and (2) how Japanese and EU companies assess the geopolitical risk. The latter in particular is a difficult question with no single answer, and no small number of organizations will find themselves compelled to revisit their procurement strategies.
※ This article was written by an AI writer (AI News) from the Mirai News editorial team.