Alibaba "Qwen 3.5" — Chinese LLM Reaches the Top, Open Weights Target Industrial Deployment
機械翻訳 / Machine-translated
Alibaba Cloud officially released its large language model "Qwen 3.5" on September 7 (Japan time). The model scored 92.4 on the mathematics reasoning benchmark "MATH-500" and 78.1 on the coding evaluation "LiveCodeBench v4," surpassing both GPT-4o (88.3 / 74.9) and Gemini 2 Pro (89.1 / 75.6) on both metrics. The 720B-parameter open weights were released the same day under the Apache 2.0 license, giving companies that have been seeking to break free from API dependency a clear alternative.
At 14:00 Japan time, Alibaba published details of Qwen 3.5 on its official blog and X account. The model is offered in three sizes:
The release notes state "support for 72 languages, with Japanese evaluation scores improved 31% over the previous generation." Breaking news also spread rapidly among Japanese-language engineering communities on X.
"I fine-tuned Qwen 3.5-72B on internal data. It's at a level where we can switch from GPT-4o. Costs dropped to less than 1/5." (Engineering account, 21,000 likes)
The Qwen series has grown rapidly over roughly three years since the first version was released in 2023. By the time Qwen 2.5 launched in early 2025, it had already earned top-tier ratings in Japanese-language and coding domains, with the expansion of open-weight versions driving ecosystem growth.
In March 2026, Alibaba Cloud announced a cumulative three-year investment of 500 billion yuan (approximately 9.6 trillion yen) in AI and cloud. The company has doubled its model development team, and the 720B model is seen as the largest output of that effort to date.
Amid ongoing US-China semiconductor restrictions, there had been speculation that Aliyun's GPU procurement faced constraints — but the benchmark results serve to demonstrate that competitive performance is achievable even under such constraints.
An API price of $0.18/M input tokens significantly undercuts Claude 5 Haiku ($0.30). Once a clear cost gap at comparable performance levels is established, adoption is expected to accelerate among cost-sensitive startups and mid-sized enterprises.
Releasing the 720B model under Apache 2.0 directly addresses large enterprises and government agencies that want to avoid vendor lock-in. Because on-premises deployment on their own cloud becomes possible, adoption in Japanese and EU markets — where data sovereignty regulations are strict — becomes realistic.
The release claims a 31% improvement over the previous generation, but the breakdown of evaluation datasets has not been disclosed. Until independent verification using standard benchmarks such as JSQuAD and JEMHopQA is compiled, these figures should not be taken at face value. Hands-on testing with internal data should come first.
As open-weight models in the 72B class become usable at lower cost than APIs, in-house development of industry-specific models is expected to accelerate. Activity is likely to emerge in sectors such as manufacturing, finance, and healthcare, where proprietary knowledge directly translates into competitive advantage.
In API mode, confirm that prompts and responses pass through Alibaba Cloud infrastructure. Running the open weights on-premises avoids this constraint, but transparency around training data remains limited. For critical use cases, it also makes sense to hold off on production deployment until third-party model audits become available.
Looking only at the numbers, the headline writes itself: "Chinese players take the top spot again." But the real significance this time lies not in the benchmark ranking — it lies in the existence of the open-weight 720B model.
Once a model with GPT-4o-class performance becomes freely available under Apache 2.0, the cost assumptions that underpinned AI adoption built on pay-per-token API dependency begin to break down. For companies that previously evaluated AI vendors solely on "performance vs. cost," a third axis now enters the picture: "self-hosted open weights vs. fully managed API."
Whether the 31% improvement in Japanese performance translates into a tangible difference in real-world use cases should become clear within the next two to three weeks as independent benchmark evaluations emerge. Three things to verify: accuracy in long-form summarization, control of honorifics and writing style, and precision in specialized terminology (medical, legal, manufacturing). Once those results are in, organizations will face a genuine operational decision: wait for a domestically developed model, or move to Qwen 3.5.
The assumption that Chinese-developed models are inferior to Western alternatives in both performance and availability no longer holds after this release. The time has come to add Chinese open-weight LLMs to your AI procurement comparison framework — ask yourself whether your organization's decision-making criteria are still stuck in 2025.
Qwen 3.5 makes a compelling case across three dimensions: performance, cost, and licensing. The next moves belong to the independent benchmarking community and to the engineering community attempting Japanese-language fine-tuning. Real-world measurement data is expected to accumulate within two to three weeks, at which point the basis for adoption decisions should solidify.
This article was written by an AI writer (AI News) from the Mirai News editorial team.