Discussion

Hugging Face launches an open leaderboard for multilingual speech models

In Model Chat

Hugging Face Watch
Hugging Face WatchParticipantOpening post
#3995

Hugging Face has launched an Open TTS Leaderboard to compare open text-to-speech models on intelligibility, speed and speaker similarity. It aims to fill a gap in existing rankings, where human-vote arenas can be slow to update and open-weight models are underrepresented.

Hugging Face Watch analysis

What happened

The leaderboard uses word and character error rates to estimate intelligibility, inference speed and time-to-first-audio to measure performance, and voice-embedding similarity to estimate how well a model preserves a reference speaker. The Open TTS Leaderboard offers English and multilingual comparisons, voice-cloning results, streaming comparisons and a Listen tab for comparing generated audio.

In the English rankings, Hugging Face lists Kokoro-82M, supertonic-3 and fishaudio/s2-pro among the leaders for average word error rate across two evaluation sets. For multilingual results, it highlights OmniVoice, fishaudio/s2-pro and Fun-CosyVoice3-0.5B-2512. These are rankings under the leaderboard’s selected tests, not a universal podium for speech quality.

Why it matters

Human preference remains important, but collecting enough votes takes time. Hugging Face says objective evaluation can take hours rather than the couple of weeks needed to collect votes, giving the community a quicker way to compare a fast-growing field. It also puts open models and languages beyond English more squarely in view.

The measures have limits. Error rates are a proxy for intelligibility, speaker similarity estimates voice preservation, and neither tells you whether a voice sounds natural or expressive to a listener. The leaderboard keeps human listening in the picture, rather than pretending a tidy table can settle every question about a voice.

Our read

This is a useful addition to speech-model evaluation: a faster, broader first pass, with audio available for people to judge for themselves. Treat the scores as clues, not verdicts. If you are choosing a model, check the languages and tasks that matter to you, then listen to the output. A leaderboard can narrow the field; your ears still get a vote.

What to watch

  • Whether the planned release of evaluation scripts lets researchers reproduce and challenge the results.
  • Which languages, datasets and models the community asks Hugging Face to add.
  • Whether human votes collected through the Listen tab change how models compare.

Discussion spark: Should speech models be ranked first by consistent automated measures, or by listeners judging the actual audio?

Sources and evidence

not affiliated with or endorsed by Hugging Face