Qwen Teases 3.8-Flash-Next, a 125B Sparse MoE
Alibaba's next Qwen release pairs a large parameter pool with a tiny active footprint, promising speed without the full compute bill.
Alibaba's Qwen team is preparing to ship Qwen3.8-Flash-Next, a mixture-of-experts model listed at 125 billion total parameters but activating only about 6 billion per token. The model is expected to land imminently on ModelScope, where its placeholder page has already surfaced ahead of the announcement.
The naming signals the design goal. "Flash" points to inference speed, and the sparse architecture backs that up: by routing each token through a small slice of its experts, the model aims to deliver the knowledge capacity of a large network while keeping the per-query compute closer to that of a much smaller dense model.
Why it matters
Sparse MoE has become the dominant strategy for teams trying to balance capability against serving cost, and Qwen has leaned into it repeatedly across its lineup. A 125B/6B split is aggressive on the efficiency side, which could make the model attractive for high-throughput deployments where latency and cost per token matter as much as raw quality.
A few things to keep in mind:
- The release is billed as covering both general text and reasoning workloads.
- It is expected under the permissive Apache 2.0 license, consistent with Qwen's open-weight track record.
- Context length and benchmark details have not yet been published.
As an announcement rather than a full launch, key specifics remain unconfirmed until weights and documentation go live. If the listed figures hold, though, Qwen3.8-Flash-Next would extend the family's push toward models that are cheap to run without giving up scale.
Sources
- Visit
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
Hacker News
More in Text / LLM
IBM's Granite 4.2 Adds Reasoning to Open LLM Line
The latest update to IBM's Apache 2.0 model family leans into structured reasoning while keeping its enterprise-friendly licensing.
Zhipu's GLM-5.3 Targets Coding at a Fraction of the Cost
The open-weight MoE model from Zhipu AI aims to match frontier closed systems on coding tasks while undercutting them on price.
Liquid AI's LFM2.5-DSpark targets faster inference
The company's latest LFM2.5 variant promises up to 3.2x faster inference without leaning on cloud-scale hardware.
0 comments
No comments yet. Be the first to weigh in.