Qwen3 Family Expands into Speech Recognition
Alibaba's Qwen team has released a new 1.7-billion-parameter model designed specifically for automatic speech recognition.
The Qwen team at Alibaba has released Qwen3-ASR-1.7B, extending its new generation of models into the domain of audio processing. This 1.7-billion-parameter model is designed for automatic speech recognition (ASR), also known as speech-to-text, and is now available on the Hugging Face Hub.
Unlike general-purpose language models, Qwen3-ASR is a specialized tool focused on a single task: accurately transcribing spoken language into written text. This makes it a foundational component for a wide range of applications, from creating meeting transcripts and video subtitles to enabling voice-activated user interfaces and accessibility tools.
The release provides another strong open-source option in a field largely defined by models like OpenAI's Whisper. With its Apache 2.0 license, Qwen3-ASR-1.7B offers a permissively licensed alternative for developers and businesses to build upon without restrictive terms. Its relatively moderate size suggests a balance between performance and computational efficiency, making it potentially suitable for a variety of hardware environments.
Sources
- Visit
Qwen/Qwen3-ASR-1.7B
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.