Cohere Releases Top-Ranked Multilingual Transcription Model
The new automatic speech recognition model from Cohere Labs sets a new benchmark on the Hugging Face Open ASR Leaderboard for multilingual performance.
Cohere has released a new model for automatic speech recognition (ASR), Cohere Transcribe, immediately claiming the top position on the Hugging Face Open ASR Leaderboard. The model demonstrates state-of-the-art performance, particularly on challenging multilingual benchmarks like FLEURS (Few-shot Learning Evaluation of Universal Representations of Speech).
Trained on a large dataset of professionally transcribed audio, the model is designed to accurately convert spoken language into text across multiple languages. This capability makes it a powerful new tool for developers building voice-enabled applications, transcription services, and other features that rely on understanding human speech.
The release of a high-performing ASR model from a major AI lab like Cohere provides a strong alternative to existing leaders in the space, such as OpenAI's Whisper. As more powerful, openly-available models for speech are released, the barrier to creating sophisticated audio-based applications continues to fall for researchers and builders.
While the model's weights are publicly accessible, it is important to note the usage restrictions. Cohere Transcribe is available under a Cohere Non-Commercial License, meaning it is intended for research and non-commercial projects rather than for deployment in production commercial applications.
Sources
- Visit
CohereLabs/cohere-transcribe-03-2026
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.