Google Releases Multimodal Gemma 4 31B Model
The new 31-billion-parameter model is an instruction-tuned, 'any-to-any' powerhouse released under a permissive Apache 2.0 license.
Google DeepMind has expanded its open model lineup with the release of Gemma 4 31B Instruct. This new model is a dense, 31-billion-parameter foundation model, marking a major update and the first entry in the Gemma 4 series.
The most significant feature of this release is its native multimodality. Described as an "any-to-any" model, Gemma 4 moves beyond text-only interactions to incorporate vision capabilities, allowing it to process and understand both text and images. This positions the model as a powerful tool for building more sophisticated, context-aware AI applications.
As an instruction-tuned model—specifically designated as an "assistant" variant—Gemma 4 31B has been fine-tuned to follow user prompts and engage in helpful dialogue. This makes it well-suited for direct integration into chatbots, creative tools, and other interactive products.
Continuing its commitment to open development, Google has released Gemma 4 under the commercially-friendly Apache 2.0 license. This allows developers and businesses to freely use, modify, and deploy the model. The official weights and model card are available on Hugging Face.
Sources
- Visit
google/gemma-4-31B-it-assistant
Hugging Face
More in Any-to-Any

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
0 comments
No comments yet. Be the first to weigh in.