Vak Conformer targets speech recognition in six Indic languages
Shunya Labs releases a Conformer-based ASR model aimed at India's underserved language landscape.
Shunya Labs has published Vak Conformer, a Conformer-based automatic speech recognition model built to transcribe six Indic languages. The model is available now on Hugging Face.
The Conformer architecture, which combines convolutional layers with self-attention, has become a common backbone for modern ASR systems because it captures both local acoustic detail and longer-range context. Applying it to Indic speech is a practical choice: these languages present rich phonetic variety and, in many cases, limited high-quality training data compared with English.
Why it matters
Speech recognition for South Asian languages remains thin relative to the size of the population that speaks them. A model tuned specifically for a group of Indic languages can lower the barrier for developers building voice interfaces, transcription tools, and accessibility features in regional markets.
- Architecture: Conformer, a widely used ASR backbone
- Coverage: six Indic languages
- Availability: open weights on Hugging Face under a custom license
The release record lists few hard specifications — parameter count, training data, and benchmark figures are not detailed in the metadata — so teams evaluating it should consult the model card directly and test against their own audio. As an initial release, it establishes a foundation the developers can build on rather than a finished, benchmarked system.
Sources
- Visit
shunyalabs/vak-conformer
Hugging Face
More in Speech → Text

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Microsoft's VibeVoice ASR Goes BitNet for CPU Speech
A BitNet-quantized speech recognition model trades GPU dependence for efficient CPU inference in English and Chinese.
CrisperWhisper 2.0 Large targets verbatim transcription
A Whisper-based ASR model that keeps every filler word and stamps timestamps to the individual word, now covering English and German.
0 comments
No comments yet. Be the first to weigh in.