MOSS-TTS: A New Multilingual Text-to-Speech Model
The new system from the OpenMOSS Team uses a novel 'delay-pattern' architecture to generate natural-sounding speech in Chinese, English, and Japanese.
The OpenMOSS Team has released MOSS-TTS, a new open-source model for generating high-quality speech from text. The system is multilingual, capable of producing audio in Chinese, English, and Japanese, making it a versatile tool for a range of voice applications.
The model's key innovation lies in its architecture. MOSS-TTS is a non-autoregressive system that uses a technique called a "delay-pattern." This approach allows it to model the rhythm and prosody of speech more effectively than some traditional methods, which can result in more natural-sounding intonation without generating audio one step at a time.
A Two-Stage System
Like many modern text-to-speech systems, MOSS-TTS operates in two stages:
- First, a text-to-spectrogram model converts the input text into a mel-spectrogram, a visual representation of the sound's frequency spectrum.
- Second, a HiFi-GAN vocoder takes this spectrogram and synthesizes it into a final audio waveform.
The complete model, along with instructions for use, is available on the Hugging Face Hub. While the weights are openly accessible, they are released under a custom license that prohibits commercial use, a key consideration for developers looking to integrate the technology.
Sources
- Visit
OpenMOSS-Team/MOSS-TTS
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.