Tencent's T1 Targets Long-Horizon Terminal Work
A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.
Tencent has introduced T1, a terminal agent built to handle the kind of multi-step command-line work that trips up general-purpose models. The system is a 122-billion-parameter mixture-of-experts model tuned with reinforcement learning specifically for long-horizon tasks, according to the research paper.
The pitch is squarely about endurance. Most coding and agentic benchmarks reward short bursts of reasoning, but real terminal work — installing dependencies, debugging failed builds, chaining shell commands across many steps — punishes models that lose the thread partway through. T1 is trained to keep its footing over longer sequences, and Tencent reports state-of-the-art results on Terminal-Bench, a benchmark designed to measure exactly that.
Why it matters
Terminal agents sit at the center of the current push toward autonomous software work, where an assistant is expected to operate a shell rather than just suggest snippets. Reinforcement learning on long-horizon rewards is one of the more promising ways to close the gap between demo-quality performance and reliable execution.
- Architecture: 122B-parameter mixture-of-experts design
- Training: reinforcement learning aimed at long-horizon terminal tasks
- Result: reported SOTA on Terminal-Bench
As with many lab releases, the details that matter most for practitioners — licensing terms, context window, and how the model behaves outside benchmark conditions — will determine how widely T1 gets adopted. For now, it stands as a concrete signal that terminal competence is becoming its own optimization target.
Sources
- Visit
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
HF Papers
More in Reasoning

IFM releases K2-Horizon, a 375B open-weight MoE
The flagship model uses a mixture-of-experts design that activates just 23 billion parameters per token, keeping inference costs in check.

Tencent Previews Hunyuan Hy4, an Apache MoE Model
The company's next-generation Hunyuan language model arrives as an early preview with a permissive license and a mixture-of-experts design.
IBM's Granite 4.2 Adds Reasoning to Open LLM Line
The latest update to IBM's Apache 2.0 model family leans into structured reasoning while keeping its enterprise-friendly licensing.
0 comments
No comments yet. Be the first to weigh in.