18 Sept 2026
Small models are winning the argument and losing the evidence
Our board carries 99 models and 30 benchmarks. Of September's 25 releases, 8 carry a comparable score and 17 do not — and the 17 are mostly…
A 27-billion-parameter model compressed to 5.9 GB of ternary weights, released under Apache 2.0 — small enough to run on a laptop.
Bonsai 2 27B is a compressed model released by PrismML on 17 September 2026, built from Qwen3.8 27B. Its point is the footprint: the weights are ternary, restricted to the values minus one, zero and plus one, with FP16 group-wise scaling, which PrismML puts at 1.76 effective bits per weight and a total of 5.9 GB — more than nine times smaller than the full-precision model it is made from. It keeps a 262,144-token context window and takes text and images in, and the weights are published free under the Apache 2.0 licence on Hugging Face and GitHub. PrismML reports that it retains 98.2% of the full-precision model's aggregate benchmark performance, against 95% for the first Bonsai 27B two months earlier; it puts the compressed model at 83.9 overall to the baseline's 85.4. The seven figures behind that are PrismML's own measurements, and each is an average over a basket of tests the company chose rather than a single benchmark result: 83.95 on a knowledge-and-reasoning basket (MMLU-Redux, GPQA Diamond and AA-LCR), 81.58 on coding (HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench), 96.57 on maths (AIME 2026, AIME 2025, GSM8K and MATH-500), 82.66 on instruction following (IFBench and IFEval), 78.59 on vision (CharXiv, A-OKVQA, OmniDocBench v1.6, RealWorldQA and OCRBench v2) and 77.57 on agentic and tool-calling work (τ²-bench and BFCLv3). Because each of those is a composite with no published methodology, and no other model on this board is measured the same way, none is stored as a score here — the numbers are recorded in this description only, and every one of them is the vendor's own. The interesting trade is visible in the same table: the compression costs most on knowledge and reasoning, where it gives up 2.71 points, and on vision, where it gives up 3.05, while instruction following comes out slightly ahead of the uncompressed baseline.
| Benchmark | Official | Community avg |
|---|---|---|
| No benchmark scores yet. Be the first to add one. | ||
“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).
Sign in to add a benchmark score for this model.
Used it? Your experience helps other buyers decide.
Write a reviewNo questions yet. Be the first to ask about Bonsai 2 27B.
Everything here links back to the same verified catalogue. Pick your next stop.