The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA3.5 Lightning
NVIDIAText / LLM

NVIDIA's Nemotron 3.5 Lightning Blends Mamba and MoE

A 30B mixture-of-experts model with just 3B active parameters aims at fast, agentic coding workloads.

Aug 1, 2026
NotableOther
Nemotron-3.5-Lightning-30B-A3B

NVIDIA has published Nemotron 3.5 Lightning 30B-A3B on Hugging Face, a text-generation and reasoning model that leans on an unusual architecture to keep inference costs down. It is a mixture-of-experts design with 30 billion total parameters but only about 3 billion active per token, released here in BF16 precision.

The standout detail is the hybrid layout: the model combines Mamba-style state-space components with a mixture-of-experts structure. That pairing is meant to deliver the throughput advantages of both approaches — Mamba's linear-scaling sequence handling and MoE's ability to route work to a small slice of the network at a time.

Built for agentic coding

NVIDIA positions the release specifically for agentic coding, the kind of multi-step, tool-using workflows where a model plans, edits, and iterates rather than answering a single prompt. A few things stand out for that use case:

  • Only 3B parameters activate per token, which lowers the compute and memory bill for long, repeated inference loops.
  • The reasoning modality suggests it is tuned to produce structured, multi-step outputs.
  • The compact active footprint makes it a candidate for teams that want capable coding assistance without frontier-scale serving costs.

Why it matters: agentic tools call models constantly, so per-token efficiency compounds fast. A sparse, hybrid model that keeps quality up while cutting active compute is exactly the trade-off developers building autonomous coding agents have been looking for.

The model ships under a custom NVIDIA license rather than a standard open license, so teams should review the terms before deploying it. Context length and benchmark figures were not specified in the release record.

Sources

  • nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters30B · MoE
Active params3B active
Size65.8 GB
PrecisionBF16
ArchitectureNemotronHForCausalLM
LicenseOTHER
Downloads530K
Likes226

Modalities

Text / LLMReasoning
2 versions — view changelog

0 comments

No comments yet. Be the first to weigh in.

More in Text / LLM

Allen Institute for AI/Text / LLM

Allen Institute Open-Sources AstaBrief Report Model

The fast report-generation model behind AI2's Asta research assistant is now available under an Apache 2.0 license.

Oct 2, 2026
Kolibri-1
Aleph Alpha/Reasoning

Aleph Alpha releases Kolibri, a sovereign reasoning model

The German AI company's open-weight mixture-of-experts model targets European needs with strong German and English reasoning.

Oct 2, 2026
GLiNER2.5-Decide
Fastino/Text / LLM

Fastino's GLiNER2.5-Decide targets lean NLP tasks

A sub-1B model bundling entity extraction, intent, sentiment and topic classification arrives on Hugging Face.

Sep 23, 2026