Hugging Face Open TTS Leaderboard
Hugging Face has introduced the Open TTS Leaderboard, a scalable evaluation framework designed to standardize the assessment of multilingual text-to-speech (TTS) and voice cloning models. By replacing slow, human-centric arena voting with objective metrics, the leaderboard reduces the time required to evaluate a model from several weeks to a few hours.
Objective Evaluation Metrics
The Open TTS Leaderboard utilizes three primary objective metrics to evaluate model performance across different dimensions:
- Intelligibility: Measured using Word Error Rate (WER) and Character Error Rate (CER). The leaderboard uses Qwen3 ASR to transcribe generated audio and compare it against the original prompt.
- Speed: Evaluated via the Inverse Real-Time Factor (RTFx) for batched offline inference and Time-to-First-Audio (TTFA) for streaming latency (batch size 1), tested on H200 GPUs and CPUs.
- Speaker Similarity: Quantified using cosine similarity (SIM) between WavLM speaker embeddings of the generated audio and a reference audio clip.
Multilingual and Voice Cloning Capabilities
The leaderboard provides specialized rankings for English and multilingual performance, utilizing datasets from Seed TTS Eval and CV3 Eval.
English Performance
Models such as hexgrad/Kokoro-82M, Supertone/supertonic-3, and fishaudio/s2-pro currently lead in English WER. The platform includes Pareto plots to visualize the trade-offs between WER, batched inference speed (RTFx), and model size.
Multilingual Performance
Because English performance is not a reliable proxy for other languages, the leaderboard allows users to toggle multilingual rankings. k2-fsa/OmniVoice, fishaudio/s2-pro, and FunAudioLLM/Fun-CosyVoice3-0.5B-2512 are identified as strong multilingual models. For character-based languages (Chinese, Japanese, and Korean), the leaderboard reports CER instead of WER.
Voice Cloning
When the "Voice cloning" toggle is enabled, the leaderboard compares models that support zero-shot voice cloning. Some models, such as bosonai/higgs-tts-3-4b and openbmb/VoxCPM2, show improved average WER when a reference audio clip is provided.
Streaming Performance and Latency
The "Streaming" tab focuses on Time-to-First-Audio (TTFA), which is critical for interactive voice agents.
- Streaming Models: TTFA is measured as the time until the first audio chunk arrives.
- Non-Streaming Models: TTFA is the time required to generate the entire utterance before playback can begin.
Testing is conducted on 50 English prompts from CV3-Eval on H200 GPUs and CPUs. kyutai/pocket-tts is highlighted as a high-performing model for streaming across both hardware configurations.
Human Preference and Community Integration
While objective metrics provide scalability, Hugging Face acknowledges that human preference remains the ultimate decider for naturalness and expressiveness. To bridge this gap, the leaderboard includes a "Listen" tab where users can compare generated outputs and provide feedback.
The project aims to be community-driven, with plans to open-source the evaluation scripts to allow for contributions via GitHub Issues and Pull Requests, similar to the Open ASR Leaderboard.