Moonshot AI Releases Trillion-Parameter Kimi-K2 Model
The new Mixture-of-Experts model brings massive scale to the open-weights community, focusing on complex reasoning and coding tasks with a 128K context window.

Moonshot AI has released the weights for Kimi-K2-Instruct, a new large language model built at a massive scale. The model is described as having a trillion parameters, placing it among the largest open-access models available to researchers and developers.
Built on a Mixture-of-Experts (MoE) architecture, Kimi-K2 is designed for efficiency at scale. The model features a 128,000-token context window, enabling it to process and reason over extensive documents and complex codebases. According to the developers, the model's strengths lie in its agentic capabilities and coding performance, suggesting an aptitude for multi-step, autonomous tasks.
A New Tool for Complex Tasks
The release of a trillion-parameter model is a significant event for the open-source AI landscape. It provides a powerful foundation for building sophisticated applications that require deep reasoning and the ability to maintain context over long interactions. Key features of Kimi-K2-Instruct include:
- Massive Scale: One trillion parameters.
- Efficient Architecture: Mixture-of-Experts (MoE).
- Large Context: 128,000-token window.
- Specialization: Strong performance in coding and agentic reasoning.
The model weights and card are now available for download from the Hugging Face Hub. The model is released under a custom license, and developers should review its specific terms of use before integrating it into their work.
Sources
- Visit
moonshotai/Kimi-K2-Instruct
Hugging Face
More in Vision-Language
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.
LLaDA-UI Brings Diffusion Decoding to GUI Agents
inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.
0 comments
No comments yet. Be the first to weigh in.