Skip to content
TrustList
WA
Product News

Which AI benchmarks labs actually cite — and where the numbers stop being comparable

Editorial

By TrustList Editorial

Two evals account for more than half of every benchmark citation across 88 models. Here's what that concentration means, and where "same benchmark name" quietly stops meaning "same test."

About Which AI benchmarks labs actually cite — and where the numbers stop being comparable

Eighty-eight models sit on TrustList's AI-model board. Only fifteen distinct benchmarks appear across all of them combined, and the citations are not spread evenly — a handful of evals show up again and again, and most appear exactly once.

What's actually cited, and how often

Benchmark Category Times cited
SWE-bench Verified Code 9
GPQA Diamond Reasoning 8
MMLU General knowledge 4
MMMU Vision 3
Terminal Bench 2.1 Agent 2
OSWorld-Verified Agent 2
DeepSWE Code 2
MATH-500, CyberGym, MMLU-Pro, NL2Repo, LMArena Elo, MATH, Artificial Analysis Index, FrontierCode 1.1 1 each

SWE-bench Verified and GPQA Diamond alone account for more than half of every citation on the board. That is not TrustList's editorial choice — it reflects which evals labs themselves keep reaching for when they announce a model.

Why two evals dominate

SWE-bench Verified measures whether a model can resolve real GitHub issues; GPQA Diamond measures graduate-level science reasoning. Between them they cover the two things a 2026 model announcement is most likely to lead with — "it can code" and "it can reason" — which is a reasonable proxy for why labs converge on citing them, whatever the model's actual specialty turns out to be in practice.

The trap in a shared benchmark name

A shared name is not a guarantee of a shared measurement. This board already has two live examples: DeepSeek quoted "DeepSWE" for one model while Google quoted "DeepSWE v1.1" for another — not obviously the same test, and their scores (62.7 and 65.3) do not belong in the same ranking without checking the version first. Separately, Zhipu's GLM-5.3 claimed first place among open-source models on "Terminal Bench 3.0" while a different model on this board cites "Terminal Bench 2.1" — again, not directly comparable numbers just because the benchmark family name matches.

What this means for reading any score on the board

A benchmark cited once is not automatically weaker evidence than one cited nine times — citation frequency measures what labs chose to market, not which eval is more rigorous. But a benchmark cited only once is also untested against how other labs interpret the same name, which is precisely where version drift hides. The safest read of any single score on this board is: check the benchmark's own definition and version on the score itself (TrustList records both), and never compare two scores on a shared name without confirming they're actually the same test.


Citation counts computed directly from TrustList's live AI-model board (/benchmarks/ai-model-board) on 31 August 2026 across all 88 models and their attached benchmark scores.