DeepSeek Announces V4.1 Flash, a Cheaper Reasoning Model
The new mixture-of-experts model is billed as more capable than V4 Pro while costing less to run, and ships under an MIT license.
DeepSeek has announced V4.1 Flash, a new text-and-reasoning model that the company positions as both cheaper and more capable than its V4 Pro tier. The news surfaced through a Hacker News discussion, where the pitch centered on a familiar DeepSeek theme: pushing the price-performance frontier rather than chasing raw scale.
The model uses a mixture-of-experts (MoE) architecture, the same approach that has defined DeepSeek's recent releases and helped it keep inference costs down by activating only a fraction of its parameters per token. DeepSeek has not yet published detailed specifications — parameter counts, active parameters, and context length remain unconfirmed at announcement.
Why it matters
DeepSeek has built its reputation on making frontier-adjacent capability affordable, and a "Flash" variant that reportedly beats the heavier "Pro" model would continue that pattern. A few reasons this release is worth watching:
- It carries an MIT license, keeping the model firmly in open-weights territory for developers and companies.
- It targets reasoning workloads, an area where cost per token matters enormously.
- Undercutting its own Pro tier signals aggressive internal competition on efficiency.
For now, the announcement is early and light on hard numbers. The claims of lower cost and higher capability will need independent benchmarks and published weights to verify, but if they hold, V4.1 Flash could reset expectations for what a budget reasoning model can deliver.
Sources
- Visit
DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
Hacker News
More in Text / LLM

OpenBMB's MiniCPM5-2B targets on-device AI
The compact 2-billion-parameter model adds long-context handling and tool-calling in a footprint small enough to run locally.

inclusionAI Tunes Ling-3.0-flash for Finance
A finance-focused variant of the Ling-3.0-flash MoE model targets financial research and agentic tool use.
RWKV7-G1j arrives as a 13.3B attention-free model
The latest RWKV7 checkpoint scales the recurrent, attention-free architecture to 13.3 billion parameters under a permissive Apache 2.0 license.
0 comments
No comments yet. Be the first to weigh in.