Boson AI logo
More articles

Higgs TTS 3 lands on Hume's Voice Replication Leaderboard

The Boson AI TeamOctober 2, 2026
Hume AI recently published the Voice Replication Leaderboard, an independent evaluation of how well eleven text-to-speech models reproduce a target speaker’s voice.
Our latest TTS model, Higgs TTS 3, placed fourth overall on speaker similarity, and scored highest on accented reference voices: narrowly ahead of VoxCPM2, and ahead of Fish Audio S2 Pro, ElevenLabs v3, and ElevenLabs Multilingual v2.

What the leaderboard measures

Hume built the evaluation around something automated metrics are bad at capturing: whether a cloned voice actually sounds like the original speaker.
Three human raters scored each clip blind, without knowing which model made it. They rated it from 1 to 5 on speaker match, audio quality, and naturalness.
A fourth score measures objective speaker similarity using TitaNet embeddings. The test set has 25 reference voices: 5 standard, 5 emotionally expressive, and 15 with a range of accents, 9 of them from non-native English speakers. Every model got the same prompts.

Where Higgs TTS 3 landed

MetricHiggs TTS 3Field position
Same Speaker (human, 1 to 5)3.834th of 11
Quality (human, 1 to 5)4.38Within 0.23 of the highest score
Objective similarity (0 to 1)0.74Tied 3rd
Accented reference voices4.28Highest in the field, 0.01 ahead of VoxCPM2
On speaker similarity, Higgs TTS 3 finished ahead of ElevenLabs Multilingual v2 and ElevenLabs v3, both Cartesia Sonic models, Qwen3-TTS, Inworld TTS-2 and Microsoft VibeVoice.
Same Speaker, all 25 reference voices
VoxCPM2
4.21
LongCat-AudioDiT 3.5B
4.06
Fish Audio S2 Pro
4.03
Higgs TTS 3
3.83
Cartesia Sonic 3.5
3.70
ElevenLabs Multilingual v2
3.68
Qwen3-TTS 1.7B
3.65
Cartesia Sonic 3.6 (beta)
3.63
Inworld TTS-2
3.62
Microsoft VibeVoice 1.5B
3.46
ElevenLabs v3
2.91
Human raters, 1 to 5: does the clone sound like the original speaker? Bars run on the full 1–5 scale.

Why the accent results really matter

Voice cloning is usually demonstrated on clean studio audio from native English speakers. Real deployments rarely look like that.
The reference voice is a sales agent in Manila, a customer-service lead in Mumbai, a receptionist in Bangkok, or a founder whose English carries the accent of the language they grew up speaking. A model that clones a broadcast-standard American voice and then flattens everyone else is not usable at that scale.
Hume’s accent group is the closest thing in the evaluation to real-world use. Topping it against ElevenLabs, Cartesia, and Inworld is the result that matters most for teams building voice agents for a global customer base.
It’s also where our team at Boson AI has put a lot of our efforts. Supporting a language properly has meant going further than adding it to a list.
When we took on Thai, we found our coverage didn’t hold up with the real language. Written Thai, for example, has essentially no punctuation, which breaks assumptions many speech systems are built around.
We brought in computer scientists who spoke Thai and rebuilt the way we handled it. That’s the kind of work it takes to make coverage hold up across accents.

Try Higgs TTS 3

Higgs TTS 3 is open-weight and available on Hugging Face (nearly 100,000 downloads in the last month), and through the Higgs TTS API.
The full leaderboard, methodology, and audio samples are on Hume’s site.
#higgs-tts
#voice-cloning
#benchmark
#leaderboard