Google DeepMind's Gemma 4 Goes Multimodal and MoE
The new open-weights family adds a mixture-of-experts design, encoder-free multimodal inputs, and an optional thinking mode.
Google DeepMind has introduced Gemma 4, the latest generation of its open-weights model family, according to the Gemma 4 Technical Report. The release marks a notable architectural shift, moving the family toward a mixture-of-experts (MoE) design while broadening its native support for text, images, and reasoning-heavy tasks.
The headline changes are structural. Gemma 4 adopts an MoE approach, which activates only a subset of parameters per token to improve efficiency relative to dense models of comparable capacity. The family is also described as encoder-free for multimodal inputs, suggesting a more unified path for handling images and text without a separate vision encoder stage.
What's new in Gemma 4
- A mixture-of-experts architecture across the family
- Encoder-free multimodal handling for text and vision
- An optional thinking mode for step-by-step reasoning
- Support for long context
The thinking mode places Gemma 4 alongside a growing set of models that expose explicit reasoning behavior, letting developers trade extra compute for stronger performance on complex problems. Combined with long-context support, that positions the family for tasks that span large documents or multi-step workflows.
Gemma 4 continues to ship under the Gemma license, keeping the weights accessible for developers and researchers who want to build on or fine-tune the models directly. Why it matters: an open MoE multimodal family with reasoning controls narrows the gap between open weights and the capabilities typically reserved for closed frontier systems, giving teams more room to self-host and customize.
Sources
- Visit
Gemma 4 Technical Report
HF Papers
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.