AI Model Rankings
An independent leaderboard — models scored on a hybrid of public benchmarks, verified reviews and community votes. Transparent by design: every model links to its full profile.
| Vote | # | Model | Score | Capability | Rating | AutomationBench | OSWorld 2.0 | OSWorld-Verified | Terminal Bench 2.1 | Terminal-Bench 4.0 | Terminal-Bench-Science 0.1 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 95 | — | — | — | — | — | — | — | — | ||
| 2 | 91 +12 vs avg | 91 | — | — | — | — | 90.8 % +7.07 | — | — | ||
| 3 | 88 +9 vs avg | 88 | — | — | — | — | 87.9 % +4.17 | — | — | ||
| 4 | 87 | — | — | — | — | — | — | — | — | ||
| 5 | 86 +7 vs avg | 88 | — | — | — | 88.3 % +1.1 | — | — | — | ||
— | 6 | Alibaba · 2.4T (95B active) params · 1,000,000 ctx · Proprietary (weights announced) | 86 +7 vs avg | 86 | — | — | — | 86.1 % −1.1 | — | — | — |
| 7 | 82 | — | — | — | — | — | — | — | — | ||
| 8 | 82 +3 vs avg | 82 | — | — | — | — | 81.6 % −2.13 | — | — | ||
| 9 | 80 +1 vs avg | 75 | — | — | — | — | 74.6 % −9.13 | — | — | ||
| 10 | 79 | — | — | — | — | — | — | — | — | ||
| 11 | 73 −6 vs avg | 73 | — | — | 72.6 % | — | — | — | — | ||
| 12 | 66 | — | — | — | — | — | — | — | — | ||
| 13 | 64 | — | — | — | — | — | — | — | — | ||
| 14 | 61 | — | — | — | — | — | — | — | — | ||
| 15 | 61 | — | — | — | — | — | — | — | — | ||
| 16 | 55 | — | — | — | — | — | — | — | — | ||
| 17 | 51 | — | — | — | — | — | — | — | — | ||
| 18 | 47 −32 vs avg | 47 | — | 31.4 % | — | — | — | 55.8 % | 52.6 % | ||
| 19 | 45 | — | — | — | — | — | — | — | — | ||
| 20 | 42 | — | — | — | — | — | — | — | — | ||
| 21 | 41 | — | — | — | — | — | — | — | — | ||
| 22 | 39 | — | — | — | — | — | — | — | — | ||
| 23 | 36 | — | — | — | — | — | — | — | — | ||
| 24 | 34 | — | — | — | — | — | — | — | — | ||
| 25 | 33 | — | — | — | — | — | — | — | — | ||
| 26 | 33 | — | — | — | — | — | — | — | — | ||
| 27 | 32 | — | — | — | — | — | — | — | — | ||
| 28 | 31 | — | — | — | — | — | — | — | — | ||
| 29 | 31 | — | — | — | — | — | — | — | — | ||
| 30 | 31 | — | — | — | — | — | — | — | — | ||
| 31 | 29 | — | — | — | — | — | — | — | — | ||
| 32 | 29 | — | — | — | — | — | — | — | — | ||
| 33 | 28 | — | — | — | — | — | — | — | — | ||
| 34 | 27 | — | — | — | — | — | — | — | — | ||
| 35 | 25 | — | — | — | — | — | — | — | — | ||
| 36 | 24 | — | — | — | — | — | — | — | — | ||
| 37 | 24 | — | — | — | — | — | — | — | — | ||
| 38 | 22 | — | — | — | — | — | — | — | — | ||
| 39 | 21 | — | — | — | — | — | — | — | — | ||
| 40 | 21 | — | — | — | — | — | — | — | — | ||
| 41 | 19 | — | — | — | — | — | — | — | — | ||
| 42 | 18 | — | — | — | — | — | — | — | — | ||
| 43 | 17 | — | — | — | — | — | — | — | — | ||
| 44 | 17 | — | — | — | — | — | — | — | — | ||
| 45 | 15 | — | — | — | — | — | — | — | — | ||
| 46 | 15 | — | — | — | — | — | — | — | — | ||
| 47 | 15 | — | — | — | — | — | — | — | — | ||
| 48 | 14 | — | — | — | — | — | — | — | — | ||
| 49 | 14 | — | — | — | — | — | — | — | — | ||
| 50 | 14 | — | — | — | — | — | — | — | — | ||
| 51 | 13 | — | — | — | — | — | — | — | — | ||
| 52 | 12 | — | — | — | — | — | — | — | — | ||
| 53 | 12 | — | — | — | — | — | — | — | — | ||
| 54 | 12 | — | — | — | — | — | — | — | — | ||
| 55 | 11 | — | — | — | — | — | — | — | — | ||
| 56 | 11 | — | — | — | — | — | — | — | — | ||
| 57 | 10 | — | — | — | — | — | — | — | — | ||
| 58 | 10 | — | — | — | — | — | — | — | — | ||
| 59 | 9 | — | — | — | — | — | — | — | — | ||
| 60 | 9 | — | — | — | — | — | — | — | — | ||
| 61 | 8 | — | — | — | — | — | — | — | — | ||
| 62 | 7 | — | — | — | — | — | — | — | — | ||
| 63 | 7 | — | — | — | — | — | — | — | — | ||
| 64 | 7 | — | — | — | — | — | — | — | — | ||
| 65 | 7 | — | — | — | — | — | — | — | — | ||
| 66 | 6 | — | — | — | — | — | — | — | — | ||
| 67 | 6 | — | — | — | — | — | — | — | — | ||
| 68 | 5 | — | — | — | — | — | — | — | — | ||
— | 69 | 0 | — | — | — | — | — | — | — | — | |
— | 70 | 0 | — | — | — | — | — | — | — | — | |
— | 71 | 0 | — | — | — | — | — | — | — | — | |
— | 72 | 0 | — | — | — | — | — | — | — | — | |
— | 73 | 0 | — | — | — | — | — | — | — | — | |
— | 74 | 0 | — | — | — | — | — | — | — | — | |
— | 75 | 0 | — | — | — | — | — | — | — | — | |
| 76 | 0 | — | — | — | — | — | — | — | — | ||
— | 77 | 0 | — | — | — | — | — | — | — | — | |
— | 78 | 0 | — | — | — | — | — | — | — | — | |
— | 79 | 0 | — | — | — | — | — | — | — | — | |
— | 80 | 0 | — | — | — | — | — | — | — | — | |
— | 81 | 0 | — | — | — | — | — | — | — | — | |
— | 82 | 0 | — | — | — | — | — | — | — | — | |
— | 83 | 0 | — | — | — | — | — | — | — | — | |
— | 84 | 0 | — | — | — | — | — | — | — | — | |
— | 85 | 0 | — | — | — | — | — | — | — | — | |
— | 86 | 0 | — | — | — | — | — | — | — | — | |
— | 87 | 0 | — | — | — | — | — | — | — | — | |
— | 88 | Tencent · 770B total, 49B active (mixture of experts) params · 1,000,000 ctx | 0 | — | — | — | — | — | — | — | — |
| 89 | 0 | — | — | — | — | — | — | — | — | ||
— | 90 | 0 | — | — | — | — | — | — | — | — | |
— | 91 | 0 | — | — | — | — | — | — | — | — | |
— | 92 | 0 | — | — | — | — | — | — | — | — | |
— | 93 | 0 | — | — | — | — | — | — | — | — | |
| 94 | Alibaba · 2.4T (95B active) params · 1,000,000 ctx · Proprietary (weights announced) | 0 | — | — | — | — | — | — | — | — | |
— | 95 | 0 | — | — | — | — | — | — | — | — | |
— | 96 | 0 | — | — | — | — | — | — | — | — |
Score = hybrid of benchmark capability (60%), review rating (25%) and community votes (15%), renormalized by available signals. Capability is the mean of normalized benchmark results; green/red deltas compare each value to the average across benchmarked models. Official (editor-verified) and community-submitted values are kept separate — community submissions only affect rankings after editor approval. Your vote, rating and submitted stats all feed this score — sign-in required, so rankings stay hard to game.
Missing a model?
Suggest one and our editors will review and add it to the board.
New AI model launches
The models that launched on TrustList this month, voted up by the community.
- 1Fugu MaxAI model
Not a model — an orchestrator that routes each task to the leanest model able to solve it, at $2/$6.
- 2Fugu Ultra v2.0AI model
The capability half of Sakana's orchestration pair — 74.3 on DeepSWE, top 2 on seven of eight benchmarks.
- 3Granite 4.2 30BAI model
Apache-2.0 reasoning with a thinking switch — 512K context, 57.0 on SWE-bench Verified.
- 4Granite 4.2 8BAI model
Apache-2.0, 8B, 47.7 on SWE-bench Verified — the size that fits where a 30B does not.
- 5GLM-5.3-FlashAI model
The "Ox Alpha" model, revealed — trained and served on 100,000 domestic Chinese chips.
- 6
- 7Hy4 PreviewAI model
Tencent's open-sourced 770B mixture-of-experts model — 49B active, more than 1M tokens of context, $0.834/$2.501.
- 8Claude Fable 5.1AI model
Anthropic's top generally available model — price unchanged, cache reads cut 75%, 52.6% on Terminal-Bench-Science.
More on TrustList
Everything here links back to the same verified catalogue. Pick your next stop.
- AI-model launchesEvery model launch, newest first.
- CompaniesAgencies, consultancies and IT service providers, ranked by verified reviews.
- ProductsSoftware and SaaS with pricing, features, integrations and alternatives.
- ArticlesGuides, comparisons and research from the editorial desk and the community.
- AwardsAnnual recognition decided by verified reviews and an independent jury.
- LaunchesNew products and releases, voted up by the community every day.
- RequestsBuyers describe what they need; vendors respond directly.
- PeopleReviewers, authors and makers with public profiles.
- ComparePut up to four listings side by side before you shortlist.