Audio8 debuts a 0.6B multilingual zero-shot TTS preview
The compact text-to-speech model promises voice cloning across languages from a footprint small enough to run without heavy hardware.
Category · audio
Open text-to-speech and voice models for natural narration, voice cloning, and real-time speech — self-hostable alternatives to closed voice APIs.
59 releases
The compact text-to-speech model promises voice cloning across languages from a footprint small enough to run without heavy hardware.
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.
A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.
An open, autoregressive text-to-speech model targets English and Spanish under a permissive Apache 2.0 license.
An ultra-small, experimental text-to-speech model arrives under Apache 2.0, aimed at running speech synthesis directly on local machines.
The new text-to-speech model offers a commercially permissive alternative for developers in a field still dominated by closed-source APIs.
The new 4-billion-parameter text-to-speech model is available for non-commercial use, promising fine-grained control over vocal delivery.
Boson AI's new text-to-speech release aims for expressive, controllable voice synthesis across multiple languages.
The new open-weights text-to-speech system targets expressive, controllable voice generation across multiple languages with built-in voice cloning.
A new text-to-speech model introduces 'delay-pattern decoding' to solve common word skipping and repetition errors in parallel generation.
The new text-to-speech system adapts the decoder-only architecture of language models like Llama to generate more natural-sounding speech.
The new Supertonic 3 model supports seven languages and is optimized for local inference with the portable ONNX format.
The new text-to-speech model uses a diffusion-transformer architecture for high-quality, expressive audio and one-shot voice cloning.
The new diffusion-based model from the OpenBMB research group supports multilingual speech, emotional control, and zero-shot voice cloning.
The new open-source model from OpenMOSS-Team generates high-quality speech in multiple languages while maintaining a remarkably small footprint.
The gaming giant behind 'PUBG' has released Raon-Speech-9B, a multimodal model for English and Korean speech recognition and synthesis.
A new open-source text-to-speech model from the k2-fsa project can replicate a voice and generate speech in multiple languages from a single short audio sample.
The new diffusion-based model handles speech, music, and general audio tasks like conversion and editing within a single, versatile framework.
The 500-million-parameter model from researcher Aratako provides a high-quality, single-speaker voice under a permissive MIT license.
The new text-to-speech model can follow natural language instructions to control tone, clone voices from short clips, and speak multiple languages.
The new model, Tada-3B-ML, is designed for fine-grained control over vocal expression across more than 10 languages.
An independent researcher has released a new English text-to-speech model under a permissive license, built on a modern generative foundation.
The new system from the OpenMOSS Team uses a novel 'delay-pattern' architecture to generate natural-sounding speech in Chinese, English, and Japanese.
The new model, SoulX-Singer, can replicate a singing voice from a short audio sample and supports both English and Chinese under a permissive license.
The new text-to-speech model is optimized for the ONNX runtime, making it a promising option for efficient, on-device audio generation.
The new 600-million-parameter Qwen3-TTS model can generate speech in multiple languages and clone voices from short audio clips.
The new 600-million-parameter model from Alibaba's Qwen team can clone voices from short audio clips for multilingual speech synthesis.
Alibaba's Qwen team has released a new text-to-speech model capable of cloning voices from just a few seconds of audio.
The new 1.7-billion-parameter text-to-speech model from Alibaba's Qwen team can generate novel voices from short audio prompts.
The new 80-million-parameter text-to-speech model adapts a powerful language model architecture for efficient, open-source audio generation.
The new 1-billion-parameter model combines a Llama 3.2 base with text-to-speech to generate more natural and nuanced audio.
The new text-to-speech model uses a hybrid diffusion and autoregressive architecture for high-quality, multilingual synthesis.
The new text-to-speech model from the audio AI company supports English, Korean, and Spanish and comes in the efficient ONNX format for deployment.
The 8-billion-parameter model from Alibaba's Qwen team understands and generates spoken responses, enabling more natural audio-first applications.
A new text-to-speech model from OpenMOSS leverages the Qwen2 architecture to generate speech in both English and Chinese.
Developer 'ekwek' has released a compact 80-million-parameter text-to-speech model, notable for its unconventional use of a Qwen3 language model architecture.
The new 500-million-parameter text-to-speech model from the Qwen team offers multilingual voice cloning and emotional control.
This new text-to-speech model can replicate a voice from just a few seconds of audio, using a novel combination of flow matching and reinforcement learning.
The new 500-million-parameter text-to-speech model from OpenBMB supports both English and Chinese and can replicate a voice from a short audio sample.
The new 500-million-parameter model is designed for generating natural, long-form speech with very low latency for interactive applications.
The new text-to-speech model focuses on performance and offers voice cloning capabilities for English under a permissive MIT license.
The French AI leader expands beyond large language models with a new, 4-billion-parameter model for generating multilingual speech.
The 2-billion-parameter text-to-speech model can clone voices from a short audio sample and is available under an Apache 2.0 license.
The new 1.7 billion-parameter model from OpenMOSS is trained on conversational data to generate natural dialogue in English and Chinese.
The new Apache 2.0 licensed model uses a Llama-based architecture to generate more natural and emotionally nuanced speech from text.
Based on the Language-Free Modeling for Multilingual Text-To-Speech (LFM2) architecture, the new model offers an efficient solution for developers.
A new 16-billion-parameter model from inclusionAI uses a Mixture-of-Experts architecture to handle a wide range of audio tasks efficiently.
The new 30B Mixture-of-Experts model from Alibaba's Qwen team can process and generate content across text, image, and audio formats.
This new instruction-tuned model from Xiaomi can handle a flexible combination of audio and text inputs and outputs, from transcription to voice synthesis.
The new 500-million-parameter model offers high-quality text-to-speech and zero-shot voice cloning under a permissive license.
The new Mixture-of-Experts model from Alibaba is fine-tuned to generate detailed, multilingual descriptions for complex audio content.
The new Apache 2.0 text-to-speech model is built on a Qwen2 architecture and optimized for local inference with GGUF support.
The new 7-billion-parameter model is designed for generating long-form, multi-speaker audio in English and Chinese under a permissive MIT license.
The new open-source model specializes in generating long-form, multi-speaker audio in both English and Mandarin, mimicking a natural podcast conversation.
The new open-source model handles both speech recognition and audio generation in a single, end-to-end architecture.
The new 1.5-billion-parameter text-to-speech model is designed to generate natural, multi-speaker audio for podcasts and other long-form content.
The new 3-billion-parameter model focuses on generating expressive, multilingual speech and is fully open for commercial use under an Apache 2.0 license.
The French AI lab's new open-source model generates streaming audio in English and French under a permissive license.
Maya Research has released a 3-billion-parameter model designed to generate natural-sounding speech in Hindi and English.