IBM Releases 2B Granite Model for Multilingual Speech
The new two-billion-parameter model offers transcription capabilities for at least five major languages under a permissive Apache 2.0 license.

IBM has entered the open-source speech recognition arena with Granite Speech 4.1, a new two-billion-parameter model. Released under the permissive Apache 2.0 license, the model is designed for automatic speech recognition (ASR), also known as speech-to-text, and is available for developers to download and integrate freely.
This release provides a strong foundation for building multilingual voice applications. The model was trained to handle transcription for several languages, broadening its utility for global development teams.
Multilingual Capabilities
While details on the full training data are pending, the model explicitly supports high-quality transcription for at least five languages:
- English
- French
- German
- Italian
- Spanish
The open availability of a capable ASR model from a major enterprise tech company like IBM is a notable development. It provides a commercially viable alternative to proprietary APIs and adds another powerful option alongside existing open-source models. Developers can access the full model weights and usage instructions on its Hugging Face repository.
Sources
- Visit
ibm-granite/granite-speech-4.1-2b
Hugging Face
More in Speech → Text
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Microsoft's VibeVoice ASR brings streaming speech-to-text
A 7-billion-parameter model targets real-time, multilingual transcription with an open release on Hugging Face.

Superwhisper's s1-mini polishes raw speech-to-text output
A compact Qwen3-based model tackles the unglamorous cleanup work that makes transcripts readable.
0 comments
No comments yet. Be the first to weigh in.