OpenBMB Releases VoxCPM for Open Voice Synthesis
The new 500-million-parameter model offers high-quality text-to-speech and zero-shot voice cloning under a permissive license.

The OpenBMB research collective has released VoxCPM-0.5B, a new open-source model for speech generation. At just 500 million parameters, it's designed to be a relatively lightweight yet capable tool for developers working with synthetic audio. The model is available under a permissive Apache 2.0 license, encouraging broad adoption.
VoxCPM is built upon the architecture of the MiniCPM model family, specifically drawing from the multimodal capabilities of MiniCPM4. By extending this foundation into the audio domain, OpenBMB provides a high-quality speech synthesis model that is both accessible and efficient, continuing the trend of powerful, specialized open models in smaller weight classes.
Zero-Shot Voice Cloning
The model's primary strength lies in its ability to perform zero-shot voice cloning. This means it can replicate a person's voice from a short audio sample without requiring any specialized fine-tuning or retraining. Its core features include:
- Bilingual text-to-speech in English and Chinese.
- Zero-shot voice cloning from brief audio clips.
- High-quality, natural-sounding audio output.
For researchers and developers interested in exploring its capabilities, the model is available for download on Hugging Face. Its open license and modest size make it a compelling option for projects requiring custom voice generation or real-time speech synthesis applications.
Sources
- Visit
openbmb/VoxCPM-0.5B
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.