The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestGoogle DeepMind4-12B-it
Google DeepMindAny-to-Any

Google Releases Gemma 4, a 12B 'Any-to-Any' Model

The new 12-billion-parameter model from Google DeepMind is designed to handle a flexible mix of data types, moving beyond traditional text and image inputs.

May 23, 2026
Major releaseGemma
Gemma 4

Google DeepMind has expanded its open-weights portfolio with the release of Gemma 4 12B Instruct, a new 12-billion-parameter model. The model's key innovation is its unified 'any-to-any' multimodal architecture, designed to handle diverse data inputs and outputs seamlessly.

Unlike traditional models that are often limited to specific input-output pairs like text-to-image, Gemma 4 is built for more flexible, generalized reasoning. According to the release details on Hugging Face, its 'any-to-any' design allows it to process combinations of modalities simultaneously, a significant step toward more capable AI systems.

Why It Matters

The arrival of Gemma 4 democratizes a sophisticated architecture previously seen in much larger, closed models. By packaging these capabilities into a relatively efficient 12B parameter model, Google enables a wider range of researchers and developers to experiment with advanced multimodal applications that require less computational overhead.

The model is available under the custom Gemma license, which includes specific terms for usage and distribution. As an instruction-tuned variant, Gemma 4 12B is optimized for direct use in conversational and task-oriented applications.

Sources

  • google/gemma-4-12B-it

    Hugging Face

    Visit
  • google-deepmind/gemma v4.0.0

    GitHub

    Visit

Get the model

Hugging FaceGitHub

Specs

Parameters12B
Size23.9 GB
PrecisionBF16
ArchitectureGemma4UnifiedForConditionalGeneration
LicenseGEMMA
Downloads2.8M
Likes1.6K

Modalities

Any-to-AnyVision-LanguageText / LLM
6 versions — view changelog

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026