MisoLabs Debuts MisoTTS, an Open Voice Model
The new text-to-speech system adapts the decoder-only architecture of language models like Llama to generate more natural-sounding speech.

A new contender has entered the open-source speech synthesis space. Startup MisoLabs has released MisoTTS, a text-to-speech (TTS) model that applies popular architectural patterns from large language models to the challenge of generating human-like audio.
Unlike many traditional TTS systems, MisoTTS uses a decoder-only transformer architecture, a design heavily inspired by models in the Llama family. This approach treats audio generation as a sequence-to-sequence task, similar to how an LLM predicts the next word in a sentence. The goal is to produce more natural and expressive speech by leveraging the same principles that have dramatically advanced text generation.
MisoTTS at a Glance
The model was trained on a foundation of public domain audiobooks and currently supports two languages. Key features of the initial release include:
- Architecture: 24-layer decoder-only transformer.
- Languages: English and Japanese.
- License: Creative Commons BY-NC-SA 4.0 (non-commercial use).
The release of MisoTTS highlights a growing trend of cross-pollination in AI research, where successful architectures from one domain are adapted to solve problems in another. While its non-commercial license limits its use in products, it provides researchers and hobbyists a new tool for exploring the intersection of language and speech. The model and code are available now on the Hugging Face Hub.
Sources
- Visit
MisoLabs/MisoTTS
Hugging Face
More in Text → Speech
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
StepFun's StepAudio 3 Gen Unifies TTS and Music
A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.
Breeze-TTS-2 Brings Open Voice Cloning to English
BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.
0 comments
No comments yet. Be the first to weigh in.