The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestQwen · AlibabaQwen3-TTS
Qwen · AlibabaText → Speech

Qwen Unveils Open Model for Custom Voice Synthesis

The new 1.7-billion-parameter text-to-speech model from Alibaba's Qwen team can generate novel voices from short audio prompts.

Jan 21, 2026
NotableApache 2.0
Qwen3-TTS-12Hz-1.7B-VoiceDesign

Alibaba's Qwen team has expanded its open-source offerings with Qwen3-TTS, a new model dedicated to high-quality speech synthesis. Released under a permissive Apache 2.0 license, this 1.7-billion-parameter system marks a significant entry into the growing field of open text-to-speech (TTS) models.

The model's standout feature is its "Voice Design" capability. Unlike traditional TTS systems that rely on a fixed set of pre-recorded voices, Qwen3-TTS can generate speech in a novel voice by analyzing a short audio prompt. This allows developers to create unique voices or clone existing ones for custom applications, a feature previously common in proprietary, API-driven systems.

Multilingual and Prompt-Driven

Qwen3-TTS is designed to be multilingual and is controlled through a combination of text and audio inputs. A user provides the text to be spoken along with a reference audio clip, and the model generates speech that matches the voice characteristics of the reference. The "12Hz" in the model's name likely refers to the sampling rate of its internal audio representation, a technique used in modern neural audio codecs to efficiently model speech.

The release of a powerful, commercially-permissive voice design model like Qwen3-TTS is a notable development for the open-source AI community. It provides a foundational tool for a wide range of applications, including personalized digital assistants, dynamic video game character dialogue, and accessibility tools, without the restrictions of closed platforms.

Sources

  • Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters1.7B
Languages10 languages
Size3.8 GB
PrecisionBF16
ArchitectureQwen3TTSForConditionalGeneration
LicenseAPACHE-2.0
Downloads283.7K
Likes418

Modalities

Text → Speech

0 comments

No comments yet. Be the first to weigh in.

More in Text → Speech

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
StepFun/Text → Speech

StepFun's StepAudio 3 Gen Unifies TTS and Music

A single discrete autoregressive model handles speech, voice design, sound effects, and music generation.

Sep 10, 2026
Breeze-TTS-2
BreezeBlue/Text → Speech

Breeze-TTS-2 Brings Open Voice Cloning to English

BreezeBlue's second-generation text-to-speech model pairs voice cloning with controllable direction, all under an open release on Hugging Face.

Aug 25, 2026