GPT-6 Astra's 99.9% on ARC-AGI-3 is two numbers, not one
EditorialBy TrustList Editorial
ARC Prize's own report gives 62.7% under the shared harness and 99.9% under OpenAI's adapter — and says neither is proof of AGI. Which one describes what you'll get depends on how you integrate.
About GPT-6 Astra's 99.9% on ARC-AGI-3 is two numbers, not one
The number that travelled with GPT-6 Astra's launch on 3 September 2026 was 99.9% on ARC-AGI-3 — the benchmark designed specifically to resist memorisation and reward genuine generalisation. Greg Brockman called the release the start of an AGI era. Headlines said "saturated."
Read the actual report from ARC Prize, the organisation that runs the benchmark, and there are two numbers, not one — and the organisation that produced them says neither means what the headline implies.
The two scores
ARC Prize evaluated Astra under two harnesses:
| Harness | Score | Cost of the run |
|---|---|---|
| Standard (shared, provider-neutral interface) | 62.7% | $26,098 |
| Provider Adapter (uses OpenAI's reasoning-state features) | 99.9% | $18,817 |
The Standard harness is the one every model is scored on the same way. The Provider Adapter harness lets a model preserve what ARC Prize calls its "opaque reasoning state" between requests — a capability OpenAI's API exposes and the shared interface does not. Under the shared interface, Astra scores 62.7%. Under the interface that plays to its architecture, 99.9%.
Both are real. Both are published. They are not the same measurement, and the 37-point gap between them is the actual finding.
What ARC Prize itself says
The organisation explicitly declines the "AGI" framing. Its report states that saturating the benchmark "would not represent 'proof of achieving AGI'," and describes Astra's result as "meaningful progress towards generalization" with caveats attached: that ARC-AGI-3 "has a tightly bounded scope and format," that its environments have "deterministic, closed-ended mechanics and goals," and that it "does not represent the complexity and open-endedness of the real world."
That is the benchmark's own author saying the benchmark was not designed to answer the question the launch coverage used it to answer.
Why the harness distinction matters to a buyer
This is not a technicality. It is the same problem TrustList's AI-model board already carries in three other places: DeepSeek quoting "DeepSWE" against Google's "DeepSWE v1.1"; GLM-5.3's first place on "Terminal Bench 3.0" next to scores on "2.1"; Gemini 3.8 Flash's 90.8% and Fable 5.1's 55.8% both being "Terminal-Bench" on different versions. A shared name is not a shared measurement.
For Astra specifically: if your integration goes through a provider-neutral gateway, a routing layer, or any interface that does not carry OpenAI's reasoning state between calls, the number that describes what you will get is closer to 62.7% than to 99.9%. If you integrate directly against OpenAI's API with the adapter's features enabled, the higher figure is the relevant one — and it costs real money per task to get there.
The rest of the launch, without the framing
Strip the AGI language and Astra is a strong, expensive, deliberately gated flagship:
- $10 input / $50 output per million tokens — 2.5× GPT-5.6 Sol. Cached input $1, batch at half price, Fast mode at double.
- 1,050,000-token context, 128K output, April 2026 knowledge cutoff.
- 72.6% on OSWorld 2.0 computer use, finishing tasks in around 40 minutes against Sol's 75.
- 97.6% on FrontierMath Tier 4, 100% on ExploitBench — both OpenAI- reported.
- First model at OpenAI's "critical" cybersecurity threshold in its own preparedness framework. The public version refuses proof-of-concept exploit requests; fuller capability is reserved for vetted defenders through the Daybreak programme, which is also where the rollout started.
How TrustList records it
On the AI-model board, Astra's ARC-AGI-3 result is stored as two separate benchmarks with the harness in the name — "ARC-AGI-3 (Standard harness)" and "ARC-AGI-3 (Provider Adapter harness)" — each with ARC Prize's report as the source and the run cost noted. Collapsing them into one "99.9%" would have been the easy thing to do. It would also have been the same category of error as publishing an unsourced benchmark score in the first place.
Sources
- ARC Prize — "OpenAI's GPT-6 Astra on ARC-AGI-3" (primary; both harness scores, costs and caveats)
- VentureBeat — "'Welcome to the AGI era': OpenAI launches GPT-6 Astra"
- CNBC — OpenAI announces rollout of GPT-6 Astra (Daybreak, cyber threshold)
- Axios — OpenAI releases GPT-6 Astra, says it may represent AGI
- DataCamp — GPT-6 Astra: features, benchmarks, pricing (pricing detail, context window)
More on TrustList
Everything here links back to the same verified catalogue. Pick your next stop.
- CompaniesAgencies, consultancies and IT service providers, ranked by verified reviews.
- ProductsSoftware and SaaS with pricing, features, integrations and alternatives.
- AwardsAnnual recognition decided by verified reviews and an independent jury.
- LaunchesNew products and releases, voted up by the community every day.
- AI ModelsBenchmark scores and community ratings for every major model.
- RequestsBuyers describe what they need; vendors respond directly.
- PeopleReviewers, authors and makers with public profiles.
- ComparePut up to four listings side by side before you shortlist.