OpenBMB Releases VoxCPM2 for Expressive TTS
The new diffusion-based model from the OpenBMB research group supports multilingual speech, emotional control, and zero-shot voice cloning.
The OpenBMB research community has released VoxCPM2, a powerful new open-source model for text-to-speech synthesis. Built on a modern diffusion-based architecture, the model aims to generate high-fidelity, expressive human speech in multiple languages.
Cloning and Control
VoxCPM2's standout feature is its ability to perform zero-shot voice cloning using just a 3-to-20 second audio sample of a target voice. This allows it to generate speech in a new voice without specific training. The model also offers fine-grained control over the output, with key capabilities including:
- Cross-lingual synthesis: Generate speech in one language using a voice from another (e.g., speaking Chinese with an English speaker's vocal characteristics).
- Emotional control: Adjust the emotional tone of the generated speech.
- Multilingual support: Primarily trained on Chinese and English.
The model uses a two-stage cascaded diffusion process. The first stage converts text into a mel-spectrogram, an acoustic representation of the audio. A second-stage vocoder then converts this spectrogram into a final audio waveform, a technique known for producing high-quality results.
VoxCPM2 represents another significant step forward for open-source generative audio, providing capabilities that rival proprietary systems. It gives researchers and developers a powerful tool for creating custom voice applications. The model is available for download on the Hugging Face Hub, though users should note its custom "OpenBMB Model License" for any usage considerations.
Sources
- Visit
openbmb/VoxCPM2
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.