inclusionAI's Ring-Zero Scales Zero-RL to a Trillion Parameters
A new mixture-of-experts model learns to reason through reinforcement learning alone, without human-annotated chains of thought.
inclusionAI has introduced Ring-Zero, a mixture-of-experts reasoning model that pushes the "zero-RL" training approach to a trillion-parameter scale. According to the accompanying paper on Hugging Face, the model develops chain-of-thought reasoning as an emergent behavior, without relying on human-annotated reasoning traces.
The zero-RL premise is straightforward but ambitious: rather than teaching a model to reason by fine-tuning it on carefully curated step-by-step examples, the approach lets reasoning behavior surface through reinforcement learning directly. Ring-Zero's contribution is demonstrating that this recipe holds up at extreme scale, where a sparse mixture-of-experts architecture activates only a fraction of its total parameters per token.
Why it matters
Human-annotated reasoning data is expensive and difficult to scale, and it can bake in the biases and limits of the annotators. If reasoning can emerge from reinforcement signals alone, the path to stronger reasoning models becomes cheaper and more repeatable.
- Released under the permissive Apache 2.0 license
- Built as a mixture-of-experts model claiming trillion-parameter total scale
- Focused on emergent chain-of-thought rather than supervised reasoning traces
Ring-Zero arrives as the open-weights community increasingly treats reinforcement learning as the primary lever for reasoning gains. The details that remain to be validated — activated parameter counts, context length, and independent benchmarks — will determine how the model stacks up against other open reasoning systems, but the framing alone signals where inclusionAI is placing its bets.
Sources
- Visit
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
HF Papers
More in Reasoning
DeepSeek Ships V4-Flash, a 304B MoE Tuned for Agents
The latest checkpoint in DeepSeek's V4 line leans into agentic workflows while keeping the permissive MIT license.
DeepSeek Refreshes V4-Flash With New 0731 Checkpoint
The MIT-licensed mixture-of-experts model returns in an updated build shipping with FP8 weights for cheaper inference.

LG AI Research debuts K-EXAONE 2.0, a 750B MoE model
The new mixture-of-experts model activates 37B parameters per token and targets English, Korean, and Spanish reasoning tasks.
0 comments
No comments yet. Be the first to weigh in.