Skip to content
TrustList
N3
AI Model

Needle 3

New· 4

An 8-29 MB automation model that runs tool calls and structured extraction on a phone, a wearable or a microcontroller, with no API bill at all.

About Needle 3

Needle 3 is an automation foundation model released by Cactus Compute on 18 September 2026, built to run entirely on the device rather than behind an API. The full network is 121M parameters and ships as a single file of 8 to 29 MB in the company's CQ2 2-bit format; every depth from 2 to 20 layers of the same network is a model that can be deployed on its own, so the same release covers a microcontroller and a laptop. Cactus lists the target hardware as phones, wearables, AR glasses, smart homes, robots, cars, Macs and PCs, game consoles, televisions and microcontrollers, and quotes up to 4,000 tokens a second of decode speed. It does three things — tool calls, structured extraction and text embeddings — and by design it does not chat: the company's founder says in public that packing general conversational capacity into a model this small is the part they deliberately gave up. The weights are Apache-2.0 and published on Hugging Face, and because the model runs locally there is no per-token price at all, which is the point of it. Cactus reports results on six benchmarks — Mobile Actions (961 rows, scored on exact calls), DroidCall (200 rows, exact calls in order), BFCL v4 (3,641 rows, AST match with no call on irrelevant input), DSTC8 (1,813 turns, field F1), SNIPS gold and SNIPS 7-way (700 rows each, field F1) — and says fine-tuning lifted every subnetwork by 18 to 36 points on DroidCall. Those numbers are published only inside a chart image, so no score is stored here; what is recorded is what each benchmark measures and on how many rows. The launch claim that it matches DeepSeek V4 Flash is contested in the open: on the vendor's own launch thread a commenter reports measuring 32.2% correct tool shapes against FunctionGemma's 90.9% on their own test, several testers report the demo choosing the wrong tool on plain requests, and others argue DeepSeek V4 Flash is the wrong model to compare a tool-calling model against in the first place. Two secondary write-ups date the release to 17 September; the 18th is used here because the vendor's own blog post, the vendor's own launch thread and the release index that carries it all agree on it.

Benchmarks & AI stats

BenchmarkOfficialCommunity avg
No benchmark scores yet. Be the first to add one.

“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).

Sign in to add a benchmark score for this model.

Trust Score

4/ 100
Trust Score: New

An earned signal from verification, reviews, awards, transparency and engagement — the vendor can't buy it.

Verification
0/100 · 20%
Reviews
0/100 · 30%
Awards
0/100 · 15%
Transparency
18/100 · 20%
Engagement
0/100 · 15%
Joining soon
Recommendations
Coming soon
Complaints
Coming soon

Updated 9/19/2026

Request a demo or quote from Needle 3

Protected by reCAPTCHA — Google Privacy Policy and Terms apply.

By sending, you agree we may share your request and contact details with the provider once you confirm your email.

Reviews

Write the first review of Needle 3

Used it? Your experience helps other buyers decide.

Write a review

Questions & answers

No questions yet. Be the first to ask about Needle 3.