Mistral Large 3 Officially Launched — Surpasses GPT-4o in Math and Coding Accuracy, Reshaping Europe's LLM Strategy
機械翻訳 / Machine-translated
機械翻訳 / Machine-translated
On August 3, 2026, Mistral AI officially released its flagship model "Mistral Large 3" via API and the Mistral Platform. The model recorded 89.3% on the MATH reasoning benchmark and 92.1% on HumanEval for coding, surpassing GPT-4o across three key metrics. This marks the first time a European-origin LLM has entered the top three of global rankings.
The announcement was made at 11:00 PM Japan Time on August 3 (4:00 PM Paris Time), with the API activated simultaneously alongside the official blog post. Pricing was set at $1.8/1M input tokens and $5.4/1M output tokens — 20 to 35% cheaper than the leading models from OpenAI and Anthropic.
The context window is 128K tokens. Function call accuracy reached 88.7% on BERCLEval (a tool-use evaluation benchmark), suggesting the model has achieved a practical level of performance for agentic use cases.
"Large 3 is hitting above 92 on HumanEval. Accuracy close to o3-mini at one-third the cost. It's worth considering a switch for agentic use cases." (Tech account, 22K followers)
Mistral AI is a French startup founded in 2023 that had raised a total of $1.2 billion in Series C funding as of end of 2025. The company has pursued an open-weight strategy while remaining mindful of the EU AI Act framework, leveraging competitive pricing through models such as Mistral 7B and Mixtral 8x22B.
Large 3 is a closed model, but the company has explicitly stated its policy of offering weights under a limited license for use on the Mistral Platform and for fine-tuning. This positions it as a middle path between fully open and enterprise-only, serving as a point of differentiation.
Since April 2026, when the European Commission formally codified "ensuring EU data sovereignty" in its procurement policy, Mistral has secured multiple public sector contracts, primarily in Germany and France. Large 3 is being launched to ride that momentum.
A HumanEval score of 92.1% represents the rate at which unit-level automated tests pass. In actual development workflows, this is considered a level capable of automating the first two steps — completion suggestion and automatic test generation — out of the three-step process of "completion suggestion → automatic test generation → review" with high accuracy. The gap with OpenAI's o3 series has now narrowed to within five percentage points.
Mistral estimates that replacing GPT-4o with Large 3 in major enterprise use cases could yield annual cost savings of over $8 million at a scale of 1 trillion tokens per month. However, total cost of ownership (TCO) factoring in differences in tool accuracy and fine-tuning costs requires individual verification.
The fact that Large 3's policy logging features are designed to align with the transparency requirements imposed on high-risk categories under the EU AI Act (Article 13) represents a practical advantage. In the high-risk AI regulations for the financial and healthcare sectors that took effect in August 2026, cases where European vendor preference is explicitly written into procurement standards are on the rise.
Looking at benchmark numbers alone, it's easy to dismiss this as yet another model claiming to surpass GPT-4o — but the true value of Large 3 lies in the combination of pricing and regulatory compliance. In European public procurement, finance, and healthcare sectors, "80-point performance, half the cost, regulatory compliance" is increasingly being preferred over "100-point performance, high cost, gray-area compliance."
For Japanese companies, there are two points of relevance. One is the possibility that AI system procurement standards for European subsidiaries will change; the other is that Mistral's pricing level could serve as a bargaining chip in price negotiations against GPT-4o. The latter is expected to become a variable that domestic SIers and cloud vendors cannot afford to ignore.
The next focal point is the timeline for porting Large 3 technology to the Mistral Codestral series. If a coding-specialized model is deployed at the same price range, the competitive landscape involving GitHub Copilot, Cursor, and Devin could shift even further.
"A European-origin model can now compete on the triple combination of performance, cost, and regulatory compliance" — that is the message this announcement has sent to the market. The pricing is at a level that could trigger a review of existing contracts, and AI procurement managers would do well to consider running a comparative evaluation this week. Which components in your stack could be replaced?
This article was written by an AI writer (AI News) from the Mirai News editorial team.