NeoHorse-1-4B tunes Qwen3.5 for agentic work
A compact 4-billion-parameter model built for tool use, coding, and multi-step reasoning arrives from TokenRhythm.

TokenRhythm has released NeoHorse-1-4B, a compact language model fine-tuned from Alibaba's Qwen3.5-4B base. At 4 billion parameters, it targets a familiar sweet spot: small enough to run on modest hardware, but tuned specifically for agentic behavior rather than general chat.
The model is aimed at three overlapping tasks — tool use, code generation, and step-by-step reasoning. That framing puts it squarely in the growing category of small models designed to act as the engine inside automated workflows, where a model calls external functions and chains multiple actions together instead of simply answering a single prompt.
Why it matters
Agentic capability has largely been the domain of larger frontier models, so the appeal of a 4B option is practical:
- Lower memory and compute requirements for local or edge deployment
- Faster inference for latency-sensitive tool-calling loops
- A permissive-enough footprint to experiment without frontier-scale infrastructure
A few caveats are worth noting. The repository lists the license simply as "other," so teams should check the terms before commercial use, and no context-length or benchmark figures accompany this initial release. As a first version from a smaller publisher building atop Qwen, NeoHorse-1-4B is best treated as a candidate to evaluate against established small agentic models rather than a proven drop-in. The full details are on its Hugging Face page.
Sources
- Visit
TokenRhythm/NeoHorse-1-4B
Hugging Face
More in Text / LLM
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
Tencent's T1 Targets Long-Horizon Terminal Work
A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.

Nex-N2.5-Pro arrives as an Apache-2.0 MoE vision model
A permissively licensed multimodal mixture-of-experts model built on a Qwen3-style MoE backbone.
0 comments
No comments yet. Be the first to weigh in.