The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIAv1
NVIDIASpeech → Text

NVIDIA's Parakeet ASR Tackles Multi-Speaker Audio

The 600-million-parameter model offers real-time speech-to-text with speaker diarization, built on the efficient FastConformer architecture.

Oct 15, 2025
NotableOther
Multitalker Parakeet Streaming 0.6B

NVIDIA has released Multitalker Parakeet Streaming 0.6B, a new open model designed to transcribe conversations with multiple participants in real time. The 600-million-parameter model addresses a common challenge for automatic speech recognition (ASR) systems: accurately capturing dialogue when more than one person is speaking.

Real-Time Diarization

The model's key capability is speaker diarization—the process of determining "who spoke when." By integrating this directly into its architecture, Parakeet can attribute transcribed text to the correct speaker as the audio is being processed. This "streaming" functionality, built on the efficient FastConformer architecture, is essential for live applications where low latency is critical.

This approach is a notable step forward for creating more useful and accurate automated transcripts. Potential applications include:

  • Live transcription and captioning for meetings and events.
  • Analyzing multi-participant audio from call centers.
  • Creating searchable records of interviews or panel discussions.

Available now on Hugging Face, the Multitalker Parakeet Streaming 0.6B model is released under the NVIDIA Open Model License Agreement. While not a traditional permissive open-source license, it allows for broad access and use of the model's weights. Its relatively compact size could enable deployment in a wide range of on-device or cloud environments.

Sources

  • nvidia/multitalker-parakeet-streaming-0.6b-v1

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters600M
LicenseOTHER
Downloads598
Likes132

Modalities

Speech → Text

0 comments

No comments yet. Be the first to weigh in.

More in Speech → Text

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
VibeVoice-ASR-Streaming-7B
Microsoft/Speech → Text

Microsoft's VibeVoice ASR brings streaming speech-to-text

A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Sep 2, 2026
s1-mini
Superwhisper/Speech → Text

Superwhisper's s1-mini polishes raw speech-to-text output

A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.

Aug 12, 2026