Boson AI Releases Higgs Audio v2 for Expressive TTS
The new 3-billion-parameter model focuses on generating expressive, multilingual speech and is fully open for commercial use under an Apache 2.0 license.

Boson AI has introduced Higgs Audio v2, a new open-source model for text-to-speech and audio generation. With three billion parameters, the model is designed to produce expressive and natural-sounding voices across multiple languages, adding a notable new entry to the competitive audio synthesis landscape.
The release is significant not just for its scale but also for its accessibility. Higgs Audio v2 is available under a permissive Apache 2.0 license, clearing the way for both research and commercial applications. This makes it a compelling alternative to proprietary APIs, offering developers a powerful foundation for building custom audio-centric features.
A Focus on Expressive Synthesis
According to Boson AI, the model specializes in "expressive voice synthesis," aiming for a higher degree of nuance and emotion in its output compared to more monotonic TTS systems. This capability is crucial for applications requiring more natural human-like speech, such as:
- Dynamic character voices in gaming and entertainment
- Engaging narration for audiobooks and podcasts
- More sophisticated and personable virtual assistants
The Higgs Audio v2 base model is now available on Hugging Face, allowing the community to begin experimenting with its capabilities and fine-tuning it for specific use cases.
Sources
- Visit
bosonai/higgs-audio-v2-generation-3B-base
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.