Meta's "Llama 4 Scout" Commercial API Goes Live — 70% Cost Reduction Reshapes the OSS LLM Competitive Landscape
機械翻訳 / Machine-translated
On August 13, 2026, Meta officially launched the commercial API (Llama API) for "Llama 4 Scout." At $0.11 per million input tokens — roughly 30% of GPT-4o's price — and equipped with a context window of up to 512k tokens and multimodal input support, this model is now available through a paid API. The development is beginning to dismantle the long-held notion that OSS models are "cheaper but inferior in performance."
Meta had already released the weights for Llama 4 Scout (17B active parameters, 109B total MoE architecture) in April 2026, but this announcement marks the official start of paid commercial availability through the Llama API.
Key specifications at the time of the official announcement are as follows:
On X, reactions from engineers in Japan spread immediately after the release.
"Scout is now available as a commercial API. At this price point with 512k multimodal support, it becomes an option that can't be ignored when selecting the next use case." (Engineering-focused account, approximately 1,400 likes)
Commercial AI API pricing entered a sharp decline in the second half of 2025, with even major closed models settling into a standard range of roughly $0.30–$1.00 per million tokens. Against this backdrop, Meta appears to have shifted to a two-pronged strategy of simultaneously releasing model weights and offering a paid API.
Cost sensitivity is also high domestically. According to the Ministry of Economy, Trade and Industry's survey on generative AI adoption trends (published July 2026), 43% of domestic companies cited "high usage costs" as the biggest barrier to AI adoption, reflecting persistent demand for low-cost, commercially usable models. While Llama 4 Scout's license imposes restrictions on companies with more than 700 million monthly active users, the vast majority of Japanese companies fall within the scope of permitted commercial use.
At $0.11/1M input tokens, this represents approximately a 71% reduction compared to GPT-4o (at $0.38). However, benchmarks such as MMLU and BIG-Bench Hard suggest that a gap with GPT-4o and Gemini 2.5 Pro still remains, making this a moment where "cost-versus-performance" tradeoff judgments are critical. A practical approach would be to start with tasks that are relatively less sensitive to accuracy, such as batch processing, summarization, and classification.
512k tokens is roughly equivalent to processing 100–200 contracts at once. This creates the potential to reduce the cost of building in-house RAG pipelines in legal, accounting, and research fields, and gives companies that were considering integrating long-document management tools a reason to revisit their investment decisions.
Llama 4 Scout remains available for self-hosting. Whether the API is more economically rational than GPU procurement costs and operational overhead depends on scale. Industries with strict data sovereignty or GDPR compliance requirements are expected to continue favoring self-hosting.
Until now, Meta's primary rationale for releasing models had been the indirect benefits to its advertising business, and it had been reluctant to monetize via API. This move into paid APIs should be read as a signal of a policy change — and makes the prospect of a commercial API for the full Maverick model increasingly realistic.
The more significant development here is the structural shift represented by "Meta starting to sell its models." A major tech company rooted in OSS has now entered the commercial API market in earnest — a space previously dominated by OpenAI, Anthropic, and Google. Taken together with Mistral Large 3 (covered previously), it's fair to conclude that the second half of 2026 marks the start of an intensifying price war among OSS-based APIs.
The implications for Japanese companies are clear. The binary choice of "GPT-4o or Gemini" has effectively broken down, with Llama 4 Scout and Mistral Large 3 now entering the comparison table. The shift in selection criteria — from "model reliability" to "cost-effectiveness by task" — is unstoppable. Procurement managers are now entering an era where the ability to independently verify the correlation between benchmark scores and their own specific tasks is a required skill.
With multimodal APIs available at the $0.11/1M token level, on-site proof-of-concept projects that were previously shelved due to cost — such as image inspection in manufacturing or product catalog processing in retail — may finally begin to move forward.
The official launch of the Llama 4 Scout API marks the moment OSS models shifted from being "tools for research and experimentation" to "active participants in commercial cost competition." The next thing to watch is the timing and pricing of a commercial API for the full Llama 4 Maverick model. In an environment where multiple low-cost APIs compete side by side, the real competitive advantage in enterprise AI will be determined less by "which model to choose" and more by the ability to design "which task gets which model."
This article was written by an AI writer (AI News) from the Mirai News editorial team.