Cohere releases Apache-licensed Arabic speech model
The Cohere Labs transcription model targets Arabic and English audio under a permissive open license.
Cohere Labs has published a new speech recognition model built for Arabic, released on Hugging Face under the permissive Apache 2.0 license. The model, labeled "Transcribe Arabic 07-2026," handles automatic speech-to-text and supports both Arabic and English audio, according to the model page.
Arabic remains an underserved language in the automatic speech recognition space. Its rich morphology, wide range of regional dialects, and the gap between spoken vernaculars and Modern Standard Arabic make transcription notably harder than for English. A model tuned specifically for Arabic, rather than one where the language is an afterthought in a massive multilingual system, is a meaningful addition for developers working in the region.
Why it matters
The Apache 2.0 license is the practical headline here. It allows commercial use, modification, and redistribution with minimal friction, which lowers the barrier for startups, researchers, and public-sector teams that want to build Arabic voice applications without licensing negotiations.
- Primary focus on Arabic transcription, with English support alongside
- Released openly under Apache 2.0
- Aimed at ASR and speech-to-text workloads
Cohere Labs, the research arm behind Cohere's open model efforts, has increasingly leaned into releasing artifacts for languages beyond English. This release fits that pattern, and its real test will come as developers benchmark it against existing multilingual systems on dialectal Arabic audio in the wild.
Sources
- Visit
CohereLabs/cohere-transcribe-arabic-07-2026
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.