The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestMicrosoft1.0
MicrosoftSpeech → Text

Microsoft's VibeVoice ASR Goes BitNet for CPU Speech

A BitNet-quantized speech recognition model trades GPU dependence for efficient CPU inference in English and Chinese.

Jul 24, 2026
NotableOther
VibeVoice ASR BitNet

Microsoft has published VibeVoice ASR BitNet, a speech recognition model built to run efficiently on CPUs rather than depending on dedicated accelerators. The release marks the debut of a BitNet-quantized variant in the VibeVoice family, targeting transcription in English and Chinese.

The pitch here is about where the model runs, not just how well. BitNet-style quantization compresses model weights aggressively—the approach is associated with extremely low-bit representations—so that inference becomes practical on commodity processors. For automatic speech recognition, that opens the door to on-device or server-side transcription without the cost and scarcity of GPUs.

Why it matters

Most capable ASR systems still assume a GPU somewhere in the pipeline. A CPU-friendly model changes the deployment calculus for anyone who needs transcription at scale or on constrained hardware.

  • Multilingual support for English and Chinese out of the box
  • Quantization aimed at reducing memory and compute footprint
  • Positioned for CPU inference rather than accelerator-bound serving

The model is released under a custom license, and Microsoft has not published parameter counts or benchmark figures alongside this initial 1.0 release. Teams evaluating it for production will want to validate accuracy against their own audio, but the direction—pushing efficient speech models onto ordinary hardware—is a meaningful one for the open-weights ecosystem.

Sources

  • microsoft/VibeVoice-ASR-BitNet

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Languagesen, zh
Size11.3 GB
PrecisionFP32
ArchitectureVibeVoiceForASRTraining
LicenseOTHER
Downloads9.5K
Likes152

Modalities

Speech → Text

0 comments

No comments yet. Be the first to weigh in.

More in Speech → Text

A.X-K2 Raon Speech 21B-A3B
KRAFTON/Any-to-Any

KRAFTON releases A.X-K2 Raon speech MoE model

The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Jul 27, 2026
CrisperWhisper 2.0 Large
Nyralabs/Speech → Text

CrisperWhisper 2.0 Large targets verbatim transcription

A Whisper-based ASR model that keeps every filler word and stamps timestamps to the individual word, now covering English and German.

Jul 15, 2026
GigaAM Multilingual
Ai Sage/Speech → Text

SberDevices releases GigaAM Multilingual ASR model

An MIT-licensed speech recognition model targeting Russian, English, and Kazakh arrives on Hugging Face.

Jul 14, 2026