Google Releases Multimodal Gemma 4 31B Model
The new 31-billion-parameter model is an instruction-tuned, 'any-to-any' powerhouse released under a permissive Apache 2.0 license.
Google DeepMind has expanded its open model lineup with the release of Gemma 4 31B Instruct. This new model is a dense, 31-billion-parameter foundation model, marking a major update and the first entry in the Gemma 4 series.
The most significant feature of this release is its native multimodality. Described as an "any-to-any" model, Gemma 4 moves beyond text-only interactions to incorporate vision capabilities, allowing it to process and understand both text and images. This positions the model as a powerful tool for building more sophisticated, context-aware AI applications.
As an instruction-tuned model—specifically designated as an "assistant" variant—Gemma 4 31B has been fine-tuned to follow user prompts and engage in helpful dialogue. This makes it well-suited for direct integration into chatbots, creative tools, and other interactive products.
Continuing its commitment to open development, Google has released Gemma 4 under the commercially-friendly Apache 2.0 license. This allows developers and businesses to freely use, modify, and deploy the model. The official weights and model card are available on Hugging Face.
Sources
- Visit
google/gemma-4-31B-it-assistant
Hugging Face
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.