Hugging Face Unifies TTS Evaluation Standards, Rankings 8,000 Models
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
What You'll Learn in This Article
On September 30, 2026, Hugging Face launched the "Open TTS Leaderboard," a platform for comparing the performance of text-to-speech (TTS) AI. TTS refers to technology that converts text into human-like speech.
Over 8,000 TTS models are currently published on the Hugging Face platform. However, evaluation methods varied widely, making it difficult to determine which models were truly superior.
This leaderboard measures all models against the same criteria and publishes the results in a ranked format. In other words, a "common ruler" for comparing AI voice quality has been created for the first time.
Traditional TTS AI evaluation relied mainly on two approaches. The first was human voting. Platforms like TTS Arena v2 and Voice Arena, for example, let users listen to multiple audio samples and vote for their favorite.
This method had its weaknesses, however. Human preferences shift with time and mood, meaning results from the same person were often inconsistent. It also took weeks for results to emerge.
The second problem was an English-centric bias. Most leaderboards measured performance in English only. Models that performed well in English often delivered lower quality in Japanese or Chinese, creating a gap between benchmark scores and real-world usability.
The Open TTS Leaderboard addresses both issues through objective numerical metrics and multilingual support. Evaluation time has also been dramatically reduced from weeks to just a few hours.
The Open TTS Leaderboard assesses models across five metrics. Let's look at what each one measures.
1. WER / CER (Intelligibility)
These metrics measure how accurately generated speech can be understood. WER stands for Word Error Rate and CER for Character Error Rate. WER is used for languages like English, while CER is used for Japanese, Chinese, and similar languages. Lower scores indicate clearer speech.
2. RTFx (Batch Processing Speed)
Short for Real-Time Factor, this indicates how many seconds it takes to generate one second of audio. An RTFx of 0.5, for instance, means one second of audio is produced in 0.5 seconds. Smaller values mean faster generation.
3. TTFA (Streaming Latency)
Short for Time To First Audio, this is the time from when audio generation begins to when the first sound is produced. For real-time conversational AI and voice assistants, a smaller value enables smoother interaction.
4. SIM (Speaker Similarity)
Short for Speaker Similarity, this measures the accuracy of voice cloning features — the ability to reproduce a specific person's voice. It is expressed as a value from 0 to 1, where 1 indicates a perfect match with the original voice.
5. Model Size
The amount of data that makes up an AI model. Smaller models are more lightweight and easier to run on smartphones and personal computers, though balancing size against performance is crucial.
These metrics are measured on high-performance hardware such as the GPU H200, as well as on standard CPUs — conditions designed to reflect real-world usage environments.
The Open TTS Leaderboard supports Japanese as well. Evaluation datasets include "Seed TTS Eval" and "CV3 Eval," enabling performance measurement across multiple languages.
The Japanese TTS market is thriving in 2026. Google's Gemini 3.8 TTS, for example, ranked first among nine Japanese models in Voice Arena's blind evaluation. Qwen-Audio-3.0-TTS, released by a research institution under Alibaba, supports 16 languages including Japanese and topped the overall rankings.
Japan-developed models are also emerging. "Irodori-TTS-v4-Large" is a Japanese-specialized model with 3.3 billion parameters, capable of zero-shot voice cloning (reproducing a voice from a small audio sample) and emotion control via emoji.
Notably, the quality gap between multilingual and Japanese-specialized models is narrowing. While Japanese-only models were previously needed to achieve adequate quality, the latest multilingual models have reached a sufficiently practical level for Japanese as well.
Voice cloning technology is already being applied in Japanese business settings. Here are three representative examples.
Streamlining Narration Production
If a narrator registers their voice with an AI, only the changed portions need to be auto-generated when a script is revised. This eliminates the need to book recording studios or conduct re-recordings, significantly cutting costs and turnaround times.
Mass-Producing eLearning Content
Training materials and educational content can be auto-generated in an instructor's voice. New lessons can be added or content updated even when the instructor is unavailable, reducing the burden on education departments.
Automating Customer Support
Integrating voice cloning into IVR (Interactive Voice Response) systems allows companies to greet customers in the voice of their representative or brand character. This enables 24/7 service while maintaining brand consistency.
As of 2026, "zero-shot cloning" technology — which creates high-accuracy voice clones from just 3 to 30 seconds of audio — has reached a practical level. Tools accessible to individuals, such as ElevenLabs' Instant Voice Cloning, are multiplying, lowering the barrier to adoption.
The arrival of the Open TTS Leaderboard is expected to bring two shifts to the voice AI market.
The first is an acceleration of development competition. With a shared evaluation standard in place, developers can now improve their models with clear, measurable goals. Competition to reach the top of the rankings will intensify, driving faster technological innovation.
The second is greater transparency in selection. Previously, companies choosing a TTS model had to go through extensive trial and error. The leaderboard now makes it possible to compare performance and cost-effectiveness at a glance, streamlining procurement decisions.
Hugging Face plans to open-source its evaluation scripts and add new datasets and metrics going forward. A system is also in place to collect community feedback, ensuring the leaderboard itself continues to evolve.
This article is a cross-post from AI Friends.