Which AI benchmarks labs actually cite — and where the numbers stop being comparable
EditorialBy TrustList Editorial
Two evals account for more than half of every benchmark citation across 88 models. Here's what that concentration means, and where "same benchmark name" quietly stops meaning "same test."
About Which AI benchmarks labs actually cite — and where the numbers stop being comparable
Eighty-eight models sit on TrustList's AI-model board. Only fifteen distinct benchmarks appear across all of them combined, and the citations are not spread evenly — a handful of evals show up again and again, and most appear exactly once.
What's actually cited, and how often
| Benchmark | Category | Times cited |
|---|---|---|
| SWE-bench Verified | Code | 9 |
| GPQA Diamond | Reasoning | 8 |
| MMLU | General knowledge | 4 |
| MMMU | Vision | 3 |
| Terminal Bench 2.1 | Agent | 2 |
| OSWorld-Verified | Agent | 2 |
| DeepSWE | Code | 2 |
| MATH-500, CyberGym, MMLU-Pro, NL2Repo, LMArena Elo, MATH, Artificial Analysis Index, FrontierCode 1.1 | — | 1 each |
SWE-bench Verified and GPQA Diamond alone account for more than half of every citation on the board. That is not TrustList's editorial choice — it reflects which evals labs themselves keep reaching for when they announce a model.
Why two evals dominate
SWE-bench Verified measures whether a model can resolve real GitHub issues; GPQA Diamond measures graduate-level science reasoning. Between them they cover the two things a 2026 model announcement is most likely to lead with — "it can code" and "it can reason" — which is a reasonable proxy for why labs converge on citing them, whatever the model's actual specialty turns out to be in practice.
The trap in a shared benchmark name
A shared name is not a guarantee of a shared measurement. This board already has two live examples: DeepSeek quoted "DeepSWE" for one model while Google quoted "DeepSWE v1.1" for another — not obviously the same test, and their scores (62.7 and 65.3) do not belong in the same ranking without checking the version first. Separately, Zhipu's GLM-5.3 claimed first place among open-source models on "Terminal Bench 3.0" while a different model on this board cites "Terminal Bench 2.1" — again, not directly comparable numbers just because the benchmark family name matches.
What this means for reading any score on the board
A benchmark cited once is not automatically weaker evidence than one cited nine times — citation frequency measures what labs chose to market, not which eval is more rigorous. But a benchmark cited only once is also untested against how other labs interpret the same name, which is precisely where version drift hides. The safest read of any single score on this board is: check the benchmark's own definition and version on the score itself (TrustList records both), and never compare two scores on a shared name without confirming they're actually the same test.
Citation counts computed directly from TrustList's live AI-model board
(/benchmarks/ai-model-board) on 31 August 2026 across all 88 models and
their attached benchmark scores.
More on TrustList
Everything here links back to the same verified catalogue. Pick your next stop.
- CompaniesAgencies, consultancies and IT service providers, ranked by verified reviews.
- ProductsSoftware and SaaS with pricing, features, integrations and alternatives.
- AwardsAnnual recognition decided by verified reviews and an independent jury.
- LaunchesNew products and releases, voted up by the community every day.
- AI ModelsBenchmark scores and community ratings for every major model.
- RequestsBuyers describe what they need; vendors respond directly.
- PeopleReviewers, authors and makers with public profiles.
- ComparePut up to four listings side by side before you shortlist.