The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestCactus Compute3
Cactus ComputeText / LLM

Cactus Needle 3: tiny on-device tool-calling models

Cactus Compute's 8–29MB models aim to run automation and tool-calling entirely on-device, rivaling far larger cloud systems.

Sep 16, 2026
NotableApache 2.0
Cactus Needle 3

Cactus Compute has released Needle 3, a family of extremely small language models built for on-device tool calling and automation. According to the model repository, the models range from just 8MB to 29MB — small enough to ship inside an app rather than call out to a server.

The pitch is that these models can handle the structured, function-calling work that powers automation flows: interpreting a request, choosing a tool, and formatting the call. Cactus claims Needle 3 can match DeepSeek V4 Flash on automation tasks, a notable assertion for models this size, though independent benchmarks aren't yet part of the record.

Why it matters

Most tool-calling today runs in the cloud, which adds latency, cost, and privacy tradeoffs. A model measured in megabytes changes the calculus:

  • It can run locally on phones and edge hardware with no network round trip.
  • Under-1B parameters keeps memory and battery demands low.
  • The Apache 2.0 license lets developers embed and modify it freely.

The headline number to watch is real-world reliability: automation is unforgiving of malformed tool calls, and sub-30MB models leave little room for error. If Cactus's claims hold up under scrutiny, Needle 3 points toward a future where routine agentic tasks happen entirely on the device in your hand.

Sources

  • Cactus-Compute/needle3

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

ArchitectureNeedleForToolCalling
LicenseAPACHE-2.0
Downloads36.1K
Likes113

Modalities

CodeText / LLM

0 comments

No comments yet. Be the first to weigh in.

More in Text / LLM

Ternary-Bonsai-2-27B
Prism Ml/Text / LLM

Ternary-Bonsai-2 packs a 27B model into 2-bit form

A ternary-quantized 27B model with hybrid attention targets on-device inference across CUDA and Metal.

Sep 16, 2026
Xing4.0-29B-A4B
XingChen AGI/Text / LLM

Xing4.0 arrives as a 29B MoE with 4B active params

inclusionAI's new text model uses a mixture-of-experts design to keep compute low while shipping under an Apache-2.0 license.

Sep 16, 2026
Agnes-3.0-Flash
Agnes AI/Vision-Language

Agnes-3.0-Flash arrives as a multimodal reasoning model

The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.

Sep 11, 2026