Skip to content
TrustList
Artificial Intelligence News

Small models are winning the argument and losing the evidence

Editorial

By TrustList Editorial

Our board carries 99 models and 30 benchmarks. Of September's 25 releases, 8 carry a comparable score and 17 do not — and the 17 are mostly the small, specialised ones the trend is celebrating.

About Small models are winning the argument and losing the evidence

Small models are winning the argument and losing the evidence

The most discussed technical story on Hacker News today is a four-billion-parameter model that writes database query plans 81% faster than Postgres does. Eight places below it sits a twenty-seven-billion-parameter model compressed into 5.9 gigabytes that keeps 98.2% of the benchmark performance of the full-precision model it came from. The comments under both are about the same thing: the era of one enormous general model is ending, and small purpose-built models are taking the work.

The direction is real. Our own launch board has recorded it all month. But there is a second thing happening at the same time, and almost nobody is talking about it, because you can only see it if you have been trying to file these releases into a comparable table.

As the models get smaller and more specialised, the evidence about them stops being comparable. Not thinner. Not more modest. Incomparable — published in a shape that cannot be set next to anything else.

Here is what that looks like counted, on our AI-model board this afternoon.

What the board actually holds

The board carries 99 models and 30 benchmark definitions. Of those 99 models, 26 carry at least one stored benchmark score. The other 73 carry none.

Narrow it to this month, where the record is freshest and every entry was checked against the vendor's own page as it was added. Twenty-five models were released in September 2026. Eight of them carry a stored score. Seventeen do not.

The eight are Claude Fable 5.1 (four scores), GPT-6 Astra (five), SWE-2 (four), Pareto 26.9 (five), Gemini 3.8 Flash (two), Fugu Ultra v2.0 (two), Gemini 3.8 Flash Cyber (one) and Atria Dawn Preview (one).

Read that list again and the pattern is hard to miss. With two exceptions it is the frontier general-purpose releases — the big, expensive, do-everything models from the largest labs. Those still publish numbers you can put in a column.

The seventeen without a score are, with a few exceptions, the other kind: the small ones, the cheap ones, the compressed ones, the specialised ones, the ones built for a single job. The models the Hacker News thread is celebrating are, on this board, the models about which least can be established.

Why each blank is blank

A missing score sounds like an oversight. These are not oversights. Each one has a reason recorded against it when the listing was written, and the reasons fall into four kinds.

A ratio instead of a number. Salesforce announced Koa on 15 September, its first CRM reasoning model, post-trained from an open 120-billion-parameter base and run entirely inside its own infrastructure. Its performance claim is that it "matches or exceeds leading model performance on CRM actions with three times fewer errors" on Salesforce's own CRM benchmark. Three times fewer than what? The comparison models are not named and no absolute figure is given. There is nothing to store: a ratio against an unnamed baseline on a private benchmark is a statement about a relationship, not a measurement.

A basket presented as a score. PrismML's Bonsai 2 27B, the compression story at number eight today, publishes seven figures and an overall 83.9 against an uncompressed baseline of 85.4. But "Coding 81.58" is an average across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench, and "Math 96.57" averages AIME 2026, AIME 2025, GSM8K and MATH-500. Filing 96.57 under MATH-500 would be false; filing it under a new benchmark called "Math" would invent a standard that exists only in one vendor's marketing. The figures are real and the compression is genuinely impressive. They still cannot be set beside anything.

A benchmark nobody else has been measured on. Alibaba released Qwen3.8-Omni-Flash today, taking text, images, audio and video across a million tokens of context at $0.15 per million input tokens. Its published figure is accuracy on OmniVideoBench rising from 63.4 to 67.8 using about 45.7% fewer tokens. No other model on this board has ever been run against OmniVideoBench. A column with one entry in it is not a comparison, it is a claim with a table drawn around it. Google's two Gemini 3.8 Live models, released on 15 September, are in exactly the same position with four benchmarks of their own.

An evaluation whose answer key is somebody else's model. TypeSafe AI's Jev, also 15 September, is the sharpest case, and to the company's credit it says so itself. Jev does not generate text at all: it returns typed answers with calibrated probabilities, priced at $0.042 per million input tokens with output free. Its headline numbers come from what TypeSafe calls workflow evaluations, which score agreement with the averaged answers of two other vendors' frontier models, on workflows TypeSafe's own team wrote. The launch post states plainly that this may carry bias. A model scored against other models, on tasks set by its own maker, produces a number that measures a resemblance.

To those four, add the quieter cases. Agnes 3.0 Flash published a benchmark table that its own model card says belongs to a different checkpoint from the served API model. Inception's Mercury 2.5 claimed a 40% increase in intelligence over its predecessor without naming a single benchmark behind the figure. Moonshot's Kimi K2.8 Preview shipped with no benchmark figures at all.

Our own thumb on the scale

Now the part that would be dishonest to leave out, and that changes how the seventeen should be read.

Some of those blanks are our decision, not the vendor's silence. We refuse to store a ratio with no absolute number. We refuse to file a basket average under the name of one benchmark inside the basket. We refuse to define a benchmark from a single vendor's figure when no other model here has been measured on it. And we refuse to conflate versions: Terminal-Bench 4.0 is kept apart from 2.1, MMLU-Pro apart from MMLU, MATH-500 apart from MATH, a run without tools apart from a verified run with them.

A directory with looser rules would show a fuller table. It would also be a table in which 96.57 sat under MATH-500 for one model and a real MATH-500 result sat there for another, and no reader could tell which was which.

So the honest statement of the finding is not "seventeen vendors published nothing." Most of them published a great deal. It is: for seventeen of September's twenty-five releases, nothing they published could be placed beside another model without misrepresenting it. That is a claim about the shape of the evidence, and it is the claim the Hacker News thread is missing while it argues about parameter counts.

It is worth saying which way the incentive runs. A vendor with a comparable number that flatters it publishes that number. When the published evidence stops being comparable, the cheapest explanation is usually not that the benchmark did not exist.

The one that published its losses

One September release went the other way, and it deserves naming because it is the counter-example that shows the rest is a choice rather than a necessity.

Unbiased AI released Pareto 26.9 on 17 September. It is not even a single set of weights: it runs several models against each other on every request and keeps the best answer, and the company says the composition can change between releases. It is exactly the sort of product that could hide behind "it's a system, not a model."

Instead it published five figures against three named competitors, on benchmarks those competitors are also measured on — and on four of the five, at least one of the models it named beats it. It ties for best on the fifth. On the same page it states that no measured task cost or composite score exists for the release, and that no independent evaluator has scored it.

Five stored scores, from the smallest and least established vendor in the month's list. The larger companies with the specialised models could have done the same and chose not to.

What to ask instead, while the table is empty

If you are buying in this category now, the comparable-numbers column is not coming back soon, so the questions have to change.

Ask what the number is measured against. Not "how did it score" but "score on what, run by whom, against which baseline, and is that baseline named". A ratio and a basket both survive the first question and fail the second.

Ask what the vendor publishes when it loses. A page that shows only wins tells you about the page, not the model. Publishing a loss costs a vendor nothing except the option of overclaiming, which is precisely why so few do it.

Run your own evaluation on your own work. This is the part the small-model trend actually makes easier, and it is the real good news in today's stories. A compressed 5.9 GB model under Apache 2.0 can be downloaded and tested against your own task this afternoon. A four-billion-parameter query planner can be measured against your own queries. When a model is small, cheap and open, the absence of a public benchmark matters far less, because the private one is within reach. The danger is the other combination: specialised, closed, expensive, and evidenced by a ratio.

Treat "purpose-built for your industry" as a claim about training data, not performance. It may well be true and it may well matter. It is not a measurement.

The seventeen, named

A count is worth nothing if you cannot check it, so here they are in release order: Claude Mythos 5.1 (1 September), Muse Spark 1.3 and Qwen3.8-Max-0902 (2 September), WeatherNext 3 (3 September), GPT Image 2.5 Sunburst, GPT Image 2.5 Flare and Mercury 2.5 (8 September), DeepSeek V4.1 Flash (10 September), Fugu Max, Kimi K2.8 Preview and Agnes 3.0 Flash (11 September), Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking, Jev and Koa (15 September), Bonsai 2 27B (17 September) and Qwen3.8-Omni-Flash (18 September).

Three of those deserve an asterisk, and pretending otherwise would weaken the argument rather than strengthen it. GPT Image 2.5 Sunburst and Flare are image generators and WeatherNext 3 is a weather model; there is no established public benchmark for those the way there is for text, so their blanks say something about the state of image and forecasting evaluation rather than about specialisation. Strip those three out and the count becomes fourteen of twenty-two, which does not change the shape of the finding.

Set against them, the eight with scores are four frontier general models, one orchestration system, one open-weights research release, one access-gated variant and one blend from an unknown vendor. Not one of the month's small, cheap or compressed models is in that column.

What this does not show

Twenty-five releases in eighteen days of one month, on one board, filed by one team applying one set of storage rules. A different directory with different rules would count differently, and that is the point of stating our rules above rather than presenting the number as neutral.

Nothing here measures whether any of these models is good. A model with no stored score may be excellent; Koa may well run circles around a general model on a service case, and Bonsai 2 27B may be the most useful thing released this month for anyone who needs to run a capable model on a laptop. The claim is narrower and harder to argue with: you cannot currently check, from what has been published, and neither can anyone else.

Nor does it measure adoption, revenue or retention. This board records what vendors chose to release and what they chose to publish about it.

One disclosure. TrustList owns listings on the launch board — our own Rutba products appear there, most recently Rutba Office on 10 September. None of them is an AI model, so none appears in any count in this article, and every figure here is unchanged if ours are removed.

And a last caveat about the framing. "Small models are losing the evidence" is our reading of a pattern in twenty-five rows. Several of those vendors would answer that a general benchmark was never the right way to judge a model built for one job, and they would have a point. The difficulty is that they have not replaced it with anything a buyer can check — and until they do, the fastest-moving part of this market is also the least inspectable.