The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIAv3
NVIDIASpeech → Text

NVIDIA Releases 600M Parakeet for Speech Recognition

The new FastConformer model uses a specialized training technique to improve transcription accuracy in noisy, real-world environments.

Aug 4, 2025
NotableCC BY 4.0
Parakeet TDT 0.6B v3

NVIDIA has released a new open model for automatic speech recognition (ASR) called Parakeet TDT 0.6B. As part of its NeMo toolkit for conversational AI, this 600-million-parameter model is designed to transcribe speech across multiple languages with high accuracy.

The model's architecture and training method are key to its performance. It uses a FastConformer encoder, which is known for its efficiency in processing audio sequences. The "TDT" in its name signifies Transducer with Denoising Training, a technique that makes the model more robust by training it to ignore noise and focus on the primary speech signal, a common challenge in real-world applications.

This release provides developers with a powerful and relatively lightweight tool for building speech-enabled products. With a permissive CC-BY-4.0 license, Parakeet can be freely used and modified for both research and commercial projects. Its 0.6-billion-parameter size makes it more accessible to deploy than the massive, multi-billion-parameter systems that often dominate ASR research.

By open-sourcing a specialized model like Parakeet, NVIDIA is contributing a significant building block to the conversational AI ecosystem. Developers interested in experimenting with the model can find the weights and usage instructions on the official Hugging Face repository.

Sources

  • nvidia/parakeet-tdt-0.6b-v3

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters600M
Languages25 languages
Size2.5 GB
PrecisionFP32
ArchitectureParakeetForTDT
LicenseCC-BY-4.0
Downloads599.1K
Likes1.1K

Modalities

Speech → Text

0 comments

No comments yet. Be the first to weigh in.

More in Speech → Text

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
VibeVoice-ASR-Streaming-7B
Microsoft/Speech → Text

Microsoft's VibeVoice ASR brings streaming speech-to-text

A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Sep 2, 2026
s1-mini
Superwhisper/Speech → Text

Superwhisper's s1-mini polishes raw speech-to-text output

A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.

Aug 12, 2026