Moonshot AI releases Kimi K3, a 2.8T-parameter MoE model
The open-weights multimodal model leans into coding and agentic tasks, extending Moonshot's Kimi line into a new scale bracket.
Moonshot AI has published Kimi K3, the latest entry in its Kimi family, releasing the weights openly on Hugging Face. The model is a mixture-of-experts (MoE) system with roughly 2.8 trillion total parameters, positioning it among the largest openly distributed models to date.
K3 is described as multimodal, handling both text and vision inputs, with an emphasis on reasoning. Moonshot points to coding and agentic workloads as areas of particular strength — the kinds of tasks that increasingly define how frontier-scale models are judged, from writing and debugging software to orchestrating multi-step tool use.
Why it matters
Openly released weights at this scale remain rare. Most models in the multi-trillion-parameter range stay behind APIs, so a downloadable MoE of this size gives researchers and companies a chance to study and self-host capabilities that are usually gated.
- Architecture: mixture-of-experts, ~2.8T total parameters
- Modalities: text and vision, with a reasoning focus
- Focus areas: coding and agentic performance
- Access: weights available on Hugging Face under a custom license
A few practical details, including context length and the exact license terms, aren't spelled out in the release record, so teams evaluating K3 for production should check the model card directly. As with earlier Kimi releases, the real test will be independent benchmarking and how the model holds up on long-horizon agentic tasks outside curated demos.
Sources
- Visit
moonshotai/Kimi-K3
Hugging Face
More in Text / LLM
Meituan Ships a Lighter, Sparser LongCat-Flash
The food-delivery giant's newest open model trims its mixture-of-experts design for more efficient inference under an MIT license.
DeepSeek Refreshes V4-Flash With New 0731 Checkpoint
The MIT-licensed mixture-of-experts model returns in an updated build shipping with FP8 weights for cheaper inference.
DeepSeek Ships V4-Flash, a 304B MoE Tuned for Agents
The latest checkpoint in DeepSeek's V4 line leans into agentic workflows while keeping the permissive MIT license.
0 comments
No comments yet. Be the first to weigh in.