The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA3.5-0.6b
NVIDIASpeech → Text

NVIDIA Releases Nemotron-3.5 Streaming ASR Model

The 600-million-parameter model uses a FastConformer architecture for real-time, multilingual speech-to-text applications.

May 15, 2026
NotableOther
Nemotron 3.5 ASR Streaming 0.6B

NVIDIA has released Nemotron-3.5 ASR Streaming, a new 600-million-parameter model specialized for automatic speech recognition. Designed for low-latency performance, the model targets applications that require real-time transcription of multilingual audio.

At its core, the model employs a FastConformer architecture paired with a Recurrent Neural Network Transducer (RNN-T) decoder. This design is particularly effective for streaming use cases, as it can process audio in small chunks as it arrives rather than waiting for an entire clip. NVIDIA notes that the model is "cache-aware," an optimization that helps maintain efficiency and speed during continuous audio processing.

This release provides developers with a powerful tool for building features like live captioning, voice command systems, and in-meeting transcription services. While not a general-purpose language model, its specialization makes it a significant addition to the open-source toolkit for speech-based AI.

The model is available on the Hugging Face Hub for download and use. It is released under the NVIDIA Open Model License Agreement, which permits distribution and the creation of derivative works.

Sources

  • nvidia/nemotron-3.5-asr-streaming-0.6b

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters600M
Languages40 languages
LicenseOTHER
Downloads748.8K
Likes1.1K

Modalities

Speech → Text

0 comments

No comments yet. Be the first to weigh in.

More in Speech → Text

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
VibeVoice-ASR-Streaming-7B
Microsoft/Speech → Text

Microsoft's VibeVoice ASR brings streaming speech-to-text

A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Sep 2, 2026
s1-mini
Superwhisper/Speech → Text

Superwhisper's s1-mini polishes raw speech-to-text output

A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.

Aug 12, 2026