Skip to content
TrustList
G6
Artificial Intelligence News

GPT-6 Astra's 99.9% on ARC-AGI-3 is two numbers, not one

Editorial

By TrustList Editorial

ARC Prize's own report gives 62.7% under the shared harness and 99.9% under OpenAI's adapter — and says neither is proof of AGI. Which one describes what you'll get depends on how you integrate.

About GPT-6 Astra's 99.9% on ARC-AGI-3 is two numbers, not one

The number that travelled with GPT-6 Astra's launch on 3 September 2026 was 99.9% on ARC-AGI-3 — the benchmark designed specifically to resist memorisation and reward genuine generalisation. Greg Brockman called the release the start of an AGI era. Headlines said "saturated."

Read the actual report from ARC Prize, the organisation that runs the benchmark, and there are two numbers, not one — and the organisation that produced them says neither means what the headline implies.

The two scores

ARC Prize evaluated Astra under two harnesses:

Harness Score Cost of the run
Standard (shared, provider-neutral interface) 62.7% $26,098
Provider Adapter (uses OpenAI's reasoning-state features) 99.9% $18,817

The Standard harness is the one every model is scored on the same way. The Provider Adapter harness lets a model preserve what ARC Prize calls its "opaque reasoning state" between requests — a capability OpenAI's API exposes and the shared interface does not. Under the shared interface, Astra scores 62.7%. Under the interface that plays to its architecture, 99.9%.

Both are real. Both are published. They are not the same measurement, and the 37-point gap between them is the actual finding.

What ARC Prize itself says

The organisation explicitly declines the "AGI" framing. Its report states that saturating the benchmark "would not represent 'proof of achieving AGI'," and describes Astra's result as "meaningful progress towards generalization" with caveats attached: that ARC-AGI-3 "has a tightly bounded scope and format," that its environments have "deterministic, closed-ended mechanics and goals," and that it "does not represent the complexity and open-endedness of the real world."

That is the benchmark's own author saying the benchmark was not designed to answer the question the launch coverage used it to answer.

Why the harness distinction matters to a buyer

This is not a technicality. It is the same problem TrustList's AI-model board already carries in three other places: DeepSeek quoting "DeepSWE" against Google's "DeepSWE v1.1"; GLM-5.3's first place on "Terminal Bench 3.0" next to scores on "2.1"; Gemini 3.8 Flash's 90.8% and Fable 5.1's 55.8% both being "Terminal-Bench" on different versions. A shared name is not a shared measurement.

For Astra specifically: if your integration goes through a provider-neutral gateway, a routing layer, or any interface that does not carry OpenAI's reasoning state between calls, the number that describes what you will get is closer to 62.7% than to 99.9%. If you integrate directly against OpenAI's API with the adapter's features enabled, the higher figure is the relevant one — and it costs real money per task to get there.

The rest of the launch, without the framing

Strip the AGI language and Astra is a strong, expensive, deliberately gated flagship:

  • $10 input / $50 output per million tokens — 2.5× GPT-5.6 Sol. Cached input $1, batch at half price, Fast mode at double.
  • 1,050,000-token context, 128K output, April 2026 knowledge cutoff.
  • 72.6% on OSWorld 2.0 computer use, finishing tasks in around 40 minutes against Sol's 75.
  • 97.6% on FrontierMath Tier 4, 100% on ExploitBench — both OpenAI- reported.
  • First model at OpenAI's "critical" cybersecurity threshold in its own preparedness framework. The public version refuses proof-of-concept exploit requests; fuller capability is reserved for vetted defenders through the Daybreak programme, which is also where the rollout started.

How TrustList records it

On the AI-model board, Astra's ARC-AGI-3 result is stored as two separate benchmarks with the harness in the name — "ARC-AGI-3 (Standard harness)" and "ARC-AGI-3 (Provider Adapter harness)" — each with ARC Prize's report as the source and the run cost noted. Collapsing them into one "99.9%" would have been the easy thing to do. It would also have been the same category of error as publishing an unsourced benchmark score in the first place.


Sources