25 Sept 2026
Will social media survive the AI change? What the evidence says
Our overnight sample had no social referrers, while Meta's crawlers made 5.9% of page requests. Usage, labelling rules and the EU AI Act…
An 8-29 MB automation model that runs tool calls and structured extraction on a phone, a wearable or a microcontroller, with no API bill at all.
Needle 3 is an automation foundation model released by Cactus Compute on 18 September 2026, built to run entirely on the device rather than behind an API. The full network is 121M parameters and ships as a single file of 8 to 29 MB in the company's CQ2 2-bit format; every depth from 2 to 20 layers of the same network is a model that can be deployed on its own, so the same release covers a microcontroller and a laptop. Cactus lists the target hardware as phones, wearables, AR glasses, smart homes, robots, cars, Macs and PCs, game consoles, televisions and microcontrollers, and quotes up to 4,000 tokens a second of decode speed. It does three things — tool calls, structured extraction and text embeddings — and by design it does not chat: the company's founder says in public that packing general conversational capacity into a model this small is the part they deliberately gave up. The weights are Apache-2.0 and published on Hugging Face, and because the model runs locally there is no per-token price at all, which is the point of it. Cactus reports results on six benchmarks — Mobile Actions (961 rows, scored on exact calls), DroidCall (200 rows, exact calls in order), BFCL v4 (3,641 rows, AST match with no call on irrelevant input), DSTC8 (1,813 turns, field F1), SNIPS gold and SNIPS 7-way (700 rows each, field F1) — and says fine-tuning lifted every subnetwork by 18 to 36 points on DroidCall. Those numbers are published only inside a chart image, so no score is stored here; what is recorded is what each benchmark measures and on how many rows. The launch claim that it matches DeepSeek V4 Flash is contested in the open: on the vendor's own launch thread a commenter reports measuring 32.2% correct tool shapes against FunctionGemma's 90.9% on their own test, several testers report the demo choosing the wrong tool on plain requests, and others argue DeepSeek V4 Flash is the wrong model to compare a tool-calling model against in the first place. Two secondary write-ups date the release to 17 September; the 18th is used here because the vendor's own blog post, the vendor's own launch thread and the release index that carries it all agree on it.
| Benchmark | Official | Community avg |
|---|---|---|
| No benchmark scores yet. Be the first to add one. | ||
“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).
Sign in to add a benchmark score for this model.
An earned signal from verification, reviews, awards, transparency and engagement — the vendor can't buy it.
Updated 9/19/2026
No questions yet. Be the first to ask about Needle 3.
Everything here links back to the same verified catalogue. Pick your next stop.