Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
BreezeBlue has released Breeze-TTS-2, an open text-to-speech model focused on English speech synthesis with voice cloning and controllable direction. The model is available now on Hugging Face under a custom license.
The headline feature is voice cloning, which lets users reproduce a target voice from reference audio rather than relying on a fixed set of built-in speakers. The model also emphasizes design and direction — a nod toward giving developers more control over how generated speech sounds, not just what it says.
Why it matters
Open TTS models remain relatively scarce compared to the flood of open text and image models, and controllable voice cloning is usually locked behind commercial APIs. A freely available option lowers the barrier for developers building assistants, accessibility tools, and media applications who want to run synthesis locally.
A few details are worth noting for anyone evaluating it:
- English is the sole supported language at launch
- The model ships under an "other" custom license, so terms should be reviewed before commercial use
- Parameter count and audio sample rate are not specified in the release
As a second-generation model, Breeze-TTS-2 signals ongoing investment in the family, though BreezeBlue keeps a low profile. Interested developers can find weights and documentation on the project's Hugging Face page.
Sources
- Visit
BreezeBlue/Breeze-TTS-2
Hugging Face
More in Text → Speech

Audio8 debuts a compact 0.1B preview TTS model
The lightweight text-to-speech model brings zero-shot voice cloning to a footprint small enough to run almost anywhere.
NVIDIA Opens Magpie TTS for Multilingual Voice Agents
The open-weights text-to-speech model targets low-latency, deployable voice agents across multiple languages.
NVIDIA opens Magpie TTS for multilingual voice agents
The company releases open weights for a low-latency, multilingual text-to-speech model aimed at real-time conversational systems.
0 comments
No comments yet. Be the first to weigh in.