Boson AI releases Higgs TTS v3, a 4B speech model
The new open-weights text-to-speech system targets expressive, controllable voice generation across multiple languages with built-in voice cloning.

Boson AI has published Higgs TTS v3, a 4-billion-parameter text-to-speech model now available on Hugging Face. The release positions the model as an expressive, controllable system for generating synthetic speech across multiple languages.
According to the model record, Higgs TTS v3 emphasizes three capabilities that increasingly define the modern TTS field:
- Expressive output, aimed at natural prosody and emotional range rather than flat narration
- Controllability, giving developers levers over how speech is rendered
- Voice cloning, allowing the model to reproduce a target voice
At 4B parameters, the model sits in a practical middle ground—large enough to handle nuanced multilingual synthesis, but small enough to be approachable for teams running their own inference rather than relying on closed APIs.
Why it matters
Text-to-speech has become one of the more competitive corners of open-weights AI, where the gap between proprietary services and freely available models has narrowed quickly. A controllable, multilingual model with voice cloning that ships with downloadable weights gives researchers and product builders a foundation they can inspect, fine-tune, and deploy without per-call costs. The model is distributed under a custom license, so teams will want to review the terms before commercial use.
Sources
- Visit
bosonai/higgs-tts-v3-4b
Hugging Face
More in Text → Speech

Audio8 debuts a 0.6B multilingual zero-shot TTS preview
The compact text-to-speech model promises voice cloning across languages from a footprint small enough to run without heavy hardware.

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

NVIDIA's Audex Unifies Audio Understanding and Speech
A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.
0 comments
No comments yet. Be the first to weigh in.