Microsoft Releases VibeVoice for Speech Transcription
The new open-source automatic speech recognition model handles multilingual transcription and speaker identification out of the box.

Microsoft has released VibeVoice-ASR, a new foundational model for automatic speech recognition. The system, now available on Hugging Face, is designed to convert spoken audio into written text across multiple languages.
Beyond simple transcription, VibeVoice's key capability is integrated speaker diarization—the ability to identify and label who is speaking and when. This feature is crucial for accurately transcribing conversations with multiple participants, such as meetings, interviews, or panel discussions, without requiring a separate post-processing step.
Why It Matters
The release adds a notable new entry into the competitive open-source audio space, which includes popular models like OpenAI's Whisper. While Microsoft has not yet published detailed performance benchmarks, VibeVoice’s built-in diarization offers a more streamlined solution for developers who would otherwise need to combine separate models for transcription and speaker identification.
Prospective users should take note of the licensing. According to the official model card, VibeVoice-ASR is being released for research purposes only. This will limit its immediate use in commercial products but provides a valuable new tool for the academic community exploring advanced speech processing systems.
Sources
- Visit
microsoft/VibeVoice-ASR
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.