18 Sept 2026
Small models are winning the argument and losing the evidence
Our board carries 99 models and 30 benchmarks. Of September's 25 releases, 8 carry a comparable score and 17 do not — and the 17 are mostly…
A blended model that runs several models on every request and keeps the best answer — 262K context, $2.50/$7.50 per million tokens.
Pareto 26.9 is the first model from Unbiased AI, released on 17 September 2026 and served through the company's own platform and through OpenRouter, where it is listed as unbiased/pareto (canonical id unbiased/pareto-20260917). It is not a single set of weights: Unbiased describes it as a blended model that runs a mix of frontier and open-source models against each other on every request and keeps the best answer, returning one response for one API call. The company distinguishes it from a router on the grounds that it does not switch models part-way through a conversation, so the prompt cache survives. It takes text and images and returns text, with a 262,144-token context window and up to 131,072 output tokens, and the model identifier is pareto. Pricing is $2.50 per million input tokens, $7.50 per million output tokens and $0.25 per million cached input tokens, which is a tenth of the input rate. Unbiased publishes five benchmark figures for this release against Claude Fable 5.1, GPT-6 Astra and DeepSeek V4.1 Flash, and publishes the tests it loses as well as the ones it wins: DeepSWE 74, Terminal-Bench 4.0 51, MMMU-Pro 78, HLE without tools 49 and ArXivMath 88. Those are the vendor's own measurements — the model card says no measured task costs or composite score have been published for this release, and no independent evaluator had published a score for it at the time of writing. The company also says the composition of the blend can change between releases, which is worth knowing before treating a score as a property of something fixed. Access is pay-as-you-go credits with each sign-up reviewed by hand. Unbiased and Pareto are built by Circuit & Chisel, a remote-first team in the US and Canada.
| Benchmark | Official | Community avg |
|---|---|---|
| ArXivMathreasoning | 88 % | — |
| DeepSWEcode | 74 % | — |
| HLE (no tools)reasoning | 49 % | — |
| MMMU-Provision | 78 % | — |
| Terminal-Bench 4.0agent | 51 % | — |
“Official” values are editor-approved and feed the ranking. “Community avg” is the mean of member submissions (shown for transparency; it never affects the ranking until an editor approves a value).
Sign in to add a benchmark score for this model.
An earned signal from verification, reviews, awards, transparency and engagement — the vendor can't buy it.
Updated 9/18/2026
Used it? Your experience helps other buyers decide.
Write a reviewNo questions yet. Be the first to ask about Pareto 26.9.
Everything here links back to the same verified catalogue. Pick your next stop.