Developers can now generate talking-head avatar videos from a still image and either an audio clip or Higgs TTS text input.

Higgs TTS 3 is built for voice chat: it speaks, not just reads. It turns model responses into expressive conversational speech across 100 languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects.

A benchmark for conversational proactivity in LLMs — noticing and acting on what the user implied but never said. 198 curated dialogues, 624 trigger points, 16 models, and a leaderboard where Recovery proves dramatically hard.
A real-time foundation model that brings human-like digital presence to customer conversations, virtual assistants, training, and interactive experiences

Today, we are publicly releasing Higgs STT 3, a state-of-the-art Speech-to-Text (STT / ASR) foundation model. It supports 94 languages with sophisticated language detection, advanced sentiment and semantic understanding, and outperforms whisper-v3-large by a large margin on key languages.

After a successful event in Toronto at the MScAC headquarters last October (200 participants), we are bringing the same energy to California with the 2026 Bay Area edition of our Higgs Audio Hackathon series in partnership with Eigen AI.