Qwen Releases Open 1.7B Custom Voice Synthesis Model
Alibaba's Qwen team has released a new text-to-speech model capable of cloning voices from just a few seconds of audio.
The Qwen team at Alibaba has released Qwen3-TTS, a new open-source text-to-speech (TTS) model. At 1.7 billion parameters, this model is designed to generate high-quality speech from text and is available under the permissive Apache 2.0 license, allowing for commercial use.
The standout feature of the new model is its ability to perform custom voice cloning. According to the release documentation, developers can use a short audio clip, typically between 3 and 10 seconds long, as a reference to synthesize speech in that specific voice. This capability opens up a wide range of applications for personalized and dynamic audio content.
Technical Details
The model, named Qwen3-TTS-12Hz-1.7B-CustomVoice, operates on a two-stage process. First, a text-to-acoustic model generates an initial audio representation from the input text and a voice embedding derived from the reference audio. Then, a vocoder converts this representation into the final audio waveform. The "12Hz" in its name refers to its tokenization rate, a technical detail related to how it processes audio information.
This release adds a powerful new tool to the growing ecosystem of open-source generative audio. By providing a capable, permissively licensed voice cloning model, the Qwen team is enabling developers to build more sophisticated and personalized voice applications, from custom assistants to accessibility tools. The model and usage instructions are available on the Hugging Face Hub.
Sources
- Visit
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.