The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • Hardware estimates
  • RSS feed
  • llms.txt
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestJetBrains2.1
JetBrainsCode

JetBrains ships Mellum 2.1, a reasoning MoE for code

The updated Mellum packs 12B total parameters but activates just 2.5B, pairing code specialization with a new thinking mode under Apache-2.0.

Sep 20, 2026
NotableApache 2.0
Mellum2.1-12B-A2.5B-Thinking

JetBrains has released Mellum 2.1, the latest version of its code-focused language model line. The new build is a mixture-of-experts design with 12B total parameters but only 2.5B active per token, and it adds a reasoning-oriented "Thinking" variant aimed at harder programming tasks.

The sparse layout is the headline here. By routing each token through a small fraction of the network, an MoE model can offer the capacity of a larger system while keeping inference costs closer to a 2.5B dense model. For a company whose core business is developer tooling, that efficiency matters: code assistance is latency-sensitive and often runs at scale across IDEs.

Why it matters

  • Permissive licensing. Mellum 2.1 ships under Apache-2.0, making it straightforward to fine-tune, self-host, or embed in commercial products.
  • Reasoning for code. The Thinking configuration signals a push toward multi-step problem solving rather than pure autocompletion.
  • Lean active footprint. At 2.5B active parameters, it is designed to be practical to run outside of heavyweight datacenter setups.

JetBrains has positioned Mellum as a purpose-built family rather than a general-purpose chatbot, and this update continues that focus on software development workloads. Teams evaluating open models for coding can find the weights and details on the Hugging Face repository.

Sources

  • JetBrains/Mellum2.1-12B-A2.5B-Thinking

    Hugging Face

    Visit
OlderMoondream shrinks Parakeet ASR for CPUsmoondream · Speech → Text · 3 weeks agoNewerAudio8-ASR-Infinite brings streaming bilingual speech recognitionEdge0 · Speech → Text · 3 weeks ago

Get the model

Hugging Face

Specs

Parameters12B · MoE
Active params2.5B active
Context window131K tokens
Size24.3 GB
PrecisionBF16
ArchitectureMellumForCausalLM
LicenseAPACHE-2.0
Downloads198
Likes78

Can you run it?

Runs on a laptop — about 8.7 GB at 4-bit.

  • BF16 (as published)

    32 GB GPU (RTX 5090) · 48 GB Mac

    25.8 GB
  • 8-bit

    16 GB GPU (RTX 4080 / 5070 Ti) · 24 GB Mac

    14.3 GB
  • 4-bit

    12 GB GPU (RTX 3060 / 4070) · 16 GB Mac

    8.7 GB
Your machine
GGUF builds
  • Mixture-of-experts: all 12.1B parameters must be resident even though only 2.5B are active per token — active count affects speed, not memory.
  • KV cache sized for 8,192 tokens of context; longer prompts need proportionally more.

Based on Hugging Face weights metadata. Assumes an 8K context and ~1 GB runtime overhead; actual needs vary. How we estimate


Modalities

CodeReasoning

The Weekly Weights

Every open release that mattered, one email a week.

0 comments

No comments yet. Be the first to weigh in.

More in Code

All Code →
MiMo-V2.6
Xiaomi/Vision-LanguageRuns on a laptop

Xiaomi distills MiMo V2.6 into a 9B model

The new MiMo-V2.6-Distill-Qwen-9B targets agentic workloads, coding, and tool use in a size that fits on modest hardware.

Sep 21, 2026
Cactus Needle 3
Cactus Compute/Text / LLM

Cactus Needle 3: tiny on-device tool-calling models

Cactus Compute's 8–29MB models aim to run automation and tool-calling entirely on-device, rivaling far larger cloud systems.

Sep 16, 2026
Tencent/ReasoningWorkstation GPU

Tencent's T1 Targets Long-Horizon Terminal Work

A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.

Sep 9, 2026