The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestMicrosoftStreaming-7B
MicrosoftSpeech → Text

Microsoft's VibeVoice ASR brings streaming speech-to-text

A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Sep 2, 2026
NotableOther
VibeVoice-ASR-Streaming-7B

Microsoft has published VibeVoice-ASR-Streaming-7B, a speech recognition model designed for streaming transcription, on Hugging Face. The 7-billion-parameter dense model handles automatic speech recognition across multiple languages, including English, Chinese, and Spanish.

The "streaming" designation is the key detail here. Rather than waiting for a full audio clip to finish before producing text, streaming ASR emits transcriptions incrementally as audio arrives. That makes the model a better fit for live captioning, meeting notes, voice interfaces, and any application where latency matters as much as accuracy.

Why it matters

Speech-to-text has become a crowded field, but most of the attention has gone to batch-oriented models. A capable open-weight streaming option from a major lab gives developers more flexibility to build low-latency, on-device or self-hosted transcription pipelines.

  • 7B parameters, dense (not a mixture-of-experts model)
  • Multilingual coverage spanning English, Chinese, Spanish, and more
  • Built specifically for streaming, real-time transcription

The model is distributed under a custom license, so teams evaluating it for production should review the terms on the model card before deploying. As an initial release in the VibeVoice ASR line, it also signals Microsoft's continued investment in open speech tooling.

Sources

  • microsoft/VibeVoice-ASR-Streaming-7B

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters7B
Languagesen, zh, es +
Size17.3 GB
PrecisionBF16
ArchitectureVibeVoiceForASRStreamingTraining
LicenseOTHER
Downloads839
Likes60

Modalities

Speech → Text

0 comments

No comments yet. Be the first to weigh in.

More in Speech → Text

s1-mini
Superwhisper/Speech → Text

Superwhisper's s1-mini polishes raw speech-to-text output

A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.

Aug 12, 2026
Vak Conformer
Shunyalabs/Speech → Text

Vak Conformer targets speech recognition in six Indic languages

Shunya Labs releases a Conformer-based ASR model aimed at India's underserved language landscape.

Aug 10, 2026
A.X-K2 Raon Speech 21B-A3B
KRAFTON/Any-to-Any

KRAFTON releases A.X-K2 Raon speech MoE model

The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Jul 27, 2026