The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestQwen · Alibaba3.8
Qwen · AlibabaText / LLM

Qwen releases 2.4T-parameter open MoE with 95B active

Alibaba's Qwen team pushes its largest sparse model yet, activating 95 billion parameters per token from a 2.4-trillion-parameter pool.

Aug 8, 2026
Major releaseOther
Qwen3.8-2.4T-A95B

Qwen, the AI group inside Alibaba, has published Qwen3.8-2.4T-A95B on Hugging Face — a text and reasoning model that stands as the team's largest open Mixture-of-Experts release to date. The headline figures are in the name: roughly 2.4 trillion total parameters, with about 95 billion active for any given token.

That sparse design is the whole point. Rather than firing every parameter on every forward pass, an MoE routes each token to a small subset of specialized "experts." The result is a model with the knowledge capacity of a very large network but the per-token compute closer to a mid-sized dense model. In practice, the 95B active count is what determines inference cost, while the 2.4T total sets the ceiling on what the model can store.

Why it matters

  • Scale meets openness. Frontier-scale MoE models have largely stayed behind closed APIs; shipping weights at this size continues Qwen's push to keep serious capability in the open ecosystem.
  • Reasoning focus. The model is tagged for both general text and reasoning, signaling it's aimed at the harder problem-solving workloads that increasingly define the top tier.
  • Deployment realism. A 2.4T-parameter checkpoint is not something most teams will run casually; expect adoption to concentrate among well-resourced labs and inference providers.

One caveat worth flagging: the release carries an "other" license rather than a standard permissive one, so teams should read the terms on the model page before building on it. As always with a fresh drop, independent benchmarks and real-world testing will determine how the raw parameter counts translate into everyday performance.

Sources

  • Qwen/Qwen3.8-2.4T-A95B

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters2400B · MoE
Active params95B active
Size4.9 TB
PrecisionBF16
ArchitectureQwen3_5MoeForCausalLM
LicenseOTHER
Downloads12.7K
Likes1.1K

Modalities

Text / LLMReasoning

0 comments

No comments yet. Be the first to weigh in.

More in Text / LLM

OpenMOSS/Vision-Language

OpenMOSS Debuts MOSS-VL for Real-Time Vision Interaction

A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.

Aug 14, 2026
DeepSeek-V4-Pro-0813
DeepSeek/Text / LLM

DeepSeek Releases V4-Pro-0813 With Open Weights

The Chinese lab pushes a higher-capability checkpoint of its V4 line to Hugging Face under a permissive MIT license.

Aug 13, 2026
DeepSeek-V4-Pro-0813
DeepSeek/Text / LLM

DeepSeek Releases V4-Pro, an MIT-Licensed MoE Model

The company's newest flagship targets reasoning and coding while keeping a permissive open-source license.

Aug 13, 2026