Zhipu AI Releases GLM-TTS for Zero-Shot Voice Cloning
This new text-to-speech model can replicate a voice from just a few seconds of audio, using a novel combination of flow matching and reinforcement learning.

Zhipu AI, the company behind the GLM family of large language models, has released GLM-TTS, a new model for text-to-speech synthesis. The system is capable of zero-shot voice cloning, meaning it can replicate a speaker's voice after hearing just a few seconds of an audio sample. It's designed to be bilingual, supporting both Chinese and English out of the box.
A New Approach to Synthesis
Instead of relying on more common diffusion techniques, GLM-TTS is built on a flow matching architecture. This approach can offer faster and more stable training compared to some alternatives. Uniquely, the model also incorporates reinforcement learning (RL) to fine-tune the output, specifically to improve the prosody—the rhythm, stress, and intonation—of the generated speech, making it sound more natural and expressive.
The model's core capability is its ability to take a 3- to 10-second audio prompt of a target voice and then generate new speech in that voice from any given text. This makes it a powerful tool for applications requiring personalized audio generation without extensive training data for each new voice.
GLM-TTS is available on Hugging Face, though it is released under a custom license from Zhipu AI that governs its use. Potential users should review its terms, as they differ from standard open-source licenses like Apache 2.0 or MIT.
Sources
- Visit
zai-org/GLM-TTS
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.