Cactus Needle 3: tiny on-device tool-calling models
Cactus Compute's 8–29MB models aim to run automation and tool-calling entirely on-device, rivaling far larger cloud systems.
Cactus Compute has released Needle 3, a family of extremely small language models built for on-device tool calling and automation. According to the model repository, the models range from just 8MB to 29MB — small enough to ship inside an app rather than call out to a server.
The pitch is that these models can handle the structured, function-calling work that powers automation flows: interpreting a request, choosing a tool, and formatting the call. Cactus claims Needle 3 can match DeepSeek V4 Flash on automation tasks, a notable assertion for models this size, though independent benchmarks aren't yet part of the record.
Why it matters
Most tool-calling today runs in the cloud, which adds latency, cost, and privacy tradeoffs. A model measured in megabytes changes the calculus:
- It can run locally on phones and edge hardware with no network round trip.
- Under-1B parameters keeps memory and battery demands low.
- The Apache 2.0 license lets developers embed and modify it freely.
The headline number to watch is real-world reliability: automation is unforgiving of malformed tool calls, and sub-30MB models leave little room for error. If Cactus's claims hold up under scrutiny, Needle 3 points toward a future where routine agentic tasks happen entirely on the device in your hand.
Sources
- Visit
Cactus-Compute/needle3
Hugging Face
More in Text / LLM
Ternary-Bonsai-2 packs a 27B model into 2-bit form
A ternary-quantized 27B model with hybrid attention targets on-device inference across CUDA and Metal.
Xing4.0 arrives as a 29B MoE with 4B active params
inclusionAI's new text model uses a mixture-of-experts design to keep compute low while shipping under an Apache-2.0 license.
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
0 comments
No comments yet. Be the first to weigh in.