SberDevices releases GigaAM Multilingual ASR model
An MIT-licensed speech recognition model targeting Russian, English, and Kazakh arrives on Hugging Face.
SberDevices has published GigaAM Multilingual, an automatic speech recognition model that transcribes across three languages: Russian, English, and Kazakh. The model is available now on Hugging Face under a permissive MIT license, which allows both research and commercial use.
The GigaAM family has been developed as part of Sber's broader speech and audio efforts, and this multilingual variant extends coverage beyond a single language. By pairing Russian with English and Kazakh — the latter a Turkic language — the release addresses a region where high-quality open ASR models remain relatively scarce.
Why it matters
Open, permissively licensed speech models are still uncommon for languages outside the English-centric mainstream. GigaAM Multilingual is notable for a few reasons:
- It ships under the MIT license, lowering barriers for developers who want to build products or fine-tune the model.
- It covers Russian and Kazakh alongside English, filling a gap for Central Asian and Russian-language deployments.
- It comes from a major regional developer with prior work in speech modeling.
The release page does not publish detailed parameter counts or benchmark figures, so teams evaluating the model will want to test it against their own audio. Still, for anyone working with Russian, English, or Kazakh speech, an openly licensed option is a welcome addition to the growing catalog of transcription tools.
Sources
- Visit
ai-sage/GigaAM-Multilingual
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.