Alibaba Releases CosyVoice 3 for Expressive TTS
The new 500-million-parameter text-to-speech model from the Qwen team offers multilingual voice cloning and emotional control.

Alibaba’s FunAudioLLM team, part of the group behind the Qwen model family, has released Fun-CosyVoice3, a 500-million-parameter foundation model for text-to-speech (TTS). The model is designed to generate highly natural, expressive, and controllable human-like speech, pushing the boundaries of open generative audio.
CosyVoice 3 stands out for its rich feature set, which brings it closer to capabilities offered by leading proprietary services. It provides a robust tool for developers working on sophisticated voice applications.
Cloning, Control, and Multilingual Support
The model's core strengths lie in its versatility and fine-grained control. Key features highlighted in the official release include:
- Multilingual and Accent Support: CosyVoice 3 handles over ten languages, including English, Chinese, Japanese, French, and Spanish, and can manage code-switching between them.
- Zero-Shot Voice Cloning: It can replicate a speaker’s voice from a mere 3-second audio clip, even performing cross-lingual cloning where the target language differs from the source clip.
- Expressive Control: The model allows for adjustments to emotion, style, rhythm, and prosody, enabling the generation of nuanced and context-aware speech.
While the model is available for commercial use, it is released under the tongyi-qianwen-license-1.0, which carries restrictions. Companies with more than 100 million monthly active users must seek a separate license from Alibaba, a detail developers should note before integrating it into large-scale products.
Sources
- Visit
FunAudioLLM/Fun-CosyVoice3-0.5B-2512
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.