# Qwen Teases 3.8-Flash-Next, a 125B Sparse MoE

> Alibaba's next Qwen release pairs a large parameter pool with a tiny active footprint, promising speed without the full compute bill.

Published by The Open Weights on Aug 26, 2026. Canonical: https://theopenweights.com/news/qwen3-8-flash-next-4f5f

## Key facts

- Company: Qwen · Alibaba
- Model: Qwen3.8-Flash-Next
- Version: 3.8-Flash-Next
- Category: Text / LLM
- Modalities: Text / LLM, Reasoning
- License: Apache 2.0 (Open weights, commercial use allowed)
- Parameters: 125B (mixture of experts)
- Active parameters: 6B
- Significance: notable
- Published: 2026-08-26
- Last verified: 2026-10-10
- Canonical URL: https://theopenweights.com/news/qwen3-8-flash-next-4f5f

Alibaba's Qwen team is preparing to ship **Qwen3.8-Flash-Next**, a mixture-of-experts model listed at 125 billion total parameters but activating only about 6 billion per token. The model is expected to land imminently on [ModelScope](https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next), where its placeholder page has already surfaced ahead of the announcement.

The naming signals the design goal. "Flash" points to inference speed, and the sparse architecture backs that up: by routing each token through a small slice of its experts, the model aims to deliver the knowledge capacity of a large network while keeping the per-query compute closer to that of a much smaller dense model.

## Why it matters

Sparse MoE has become the dominant strategy for teams trying to balance capability against serving cost, and Qwen has leaned into it repeatedly across its lineup. A 125B/6B split is aggressive on the efficiency side, which could make the model attractive for high-throughput deployments where latency and cost per token matter as much as raw quality.

A few things to keep in mind:

- The release is billed as covering both general text and reasoning workloads.
- It is expected under the permissive **Apache 2.0** license, consistent with Qwen's open-weight track record.
- Context length and benchmark details have not yet been published.

As an announcement rather than a full launch, key specifics remain unconfirmed until weights and documentation go live. If the listed figures hold, though, Qwen3.8-Flash-Next would extend the family's push toward models that are cheap to run without giving up scale.

## Get the model

- [Hacker News](https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next)
- [Hugging Face](https://huggingface.co/Qwen/Qwen3.8-Flash-Next)

## Sources

- [Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)](https://modelscope.cn/models/Qwen/Qwen3.8-Flash-Next) — Hacker News, Aug 25, 2026
- [Qwen/Qwen3.8-Flash-Next](https://huggingface.co/Qwen/Qwen3.8-Flash-Next) — Hugging Face, Aug 24, 2026

---
Source: The Open Weights (https://theopenweights.com/). Aggregated and written by Claude, curated by humans. Cite as: "Qwen Teases 3.8-Flash-Next, a 125B Sparse MoE", The Open Weights, Aug 26, 2026, https://theopenweights.com/news/qwen3-8-flash-next-4f5f