Nari Labs Releases Dia2-2B, an Open Voice Cloning Model
The 2-billion-parameter text-to-speech model can clone voices from a short audio sample and is available under an Apache 2.0 license.
Nari Labs has introduced Dia2-2B, a powerful new open-source model for text-to-speech (TTS) applications. The 2-billion-parameter model is designed for high-fidelity audio generation and is released under the permissive Apache 2.0 license, allowing for broad commercial and research use.
The model's primary capability is zero-shot voice cloning. It can analyze a brief audio sample to capture the unique acoustic properties of a speaker—including timbre, rhythm, and prosody—and then generate new speech in that voice from any given text. This allows for the creation of dynamic, custom voice outputs without needing to train a new model for each speaker.
Technical Foundations
Dia2-2B is a diffusion-based model, a technique known for producing high-quality generative results. It was trained on a substantial dataset of over 200,000 hours of English speech sourced from public domain audiobooks. While building on foundational concepts from the Bark model, Dia2 features a distinct architecture and was trained on a completely new dataset.
This release provides developers with a strong, openly available tool for creating sophisticated audio applications. As an alternative to proprietary TTS and voice cloning APIs, Dia2-2B enables a new class of customizable products, from personalized digital assistants to dynamic content creation tools. The model is available for download and use from the Nari Labs Hugging Face repository.
Sources
- Visit
nari-labs/Dia2-2B
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.