Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
BreezeBlue has released Breeze-TTS-2, an open text-to-speech model focused on English speech synthesis with voice cloning and controllable direction. The model is available now on Hugging Face under a custom license.
The headline feature is voice cloning, which lets users reproduce a target voice from reference audio rather than relying on a fixed set of built-in speakers. The model also emphasizes design and direction — a nod toward giving developers more control over how generated speech sounds, not just what it says.
Why it matters
Open TTS models remain relatively scarce compared to the flood of open text and image models, and controllable voice cloning is usually locked behind commercial APIs. A freely available option lowers the barrier for developers building assistants, accessibility tools, and media applications who want to run synthesis locally.
A few details are worth noting for anyone evaluating it:
- English is the sole supported language at launch
- The model ships under an "other" custom license, so terms should be reviewed before commercial use
- Parameter count and audio sample rate are not specified in the release
As a second-generation model, Breeze-TTS-2 signals ongoing investment in the family, though BreezeBlue keeps a low profile. Interested developers can find weights and documentation on the project's Hugging Face page.
Sources
- Visit
BreezeBlue/Breeze-TTS-2
Hugging Face
More in Text → Speech
All Text → Speech →Nari Labs Ships Qwen3-Based TTS and ASR Models
The startup pairs speech synthesis and recognition built on Qwen3, pitching accuracy, low latency, and lower cost.
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
0 comments
No comments yet. Be the first to weigh in.