GPT-6 Astra scored 99.9% on ARC-AGI-3. Under ARC's standard harness: 62.7%
Same model, same benchmark, two verified scores: 99.9% under OpenAI's state-preserving harness, 62.7% under ARC Prize's neutral one. Inside the 37-point gap.

Same model, same benchmark, two verified scores 37 points apart. ARC Prize published its own evaluation of GPT-6 Astra on September 3: 62.7% on its Standard harness for about $26,000, versus 98.6% to 99.9% on a Provider Adapter harness that preserves private reasoning state between requests. OpenAI's launch materials led with the higher number and the word "saturates." Reddit spent September 3–4 arguing over whether that was measurement or misleading reporting; "technically true metric reporting being used to deliberately mislead," one widely-shared post called it. Both sides describe something real; that is why the argument matters.
One model, two report cards
ARC Prize ran Astra on both harnesses at every reasoning effort level on the Semi-Private set — ARC-measured scores, not vendor-reported:
| Reasoning effort | Standard harness | Provider Adapter harness |
|---|---|---|
| max | 62.7% — $26,098 | 98.6% — $17,332 |
| xhigh | 59.3% — $37,317 | 98.4% — $18,147 |
| high | 54.8% — $40,705 | 99.9% — $18,817 |
| medium | 38.6% — $48,090 | 98.4% — $19,285 |
| low | 17.5% — $38,166 | 98.0% — $21,298 |
| none | 35.2% — $49,791 | 96.7% — $23,457 |
The columns are not two runs of the same experiment. ARC's Standard harness is deliberately minimal and provider-neutral: after each action the model's private reasoning is discarded, and only the notes it chose to write survive, carried forward in visible form. The Provider Adapter keeps the model's opaque reasoning state alive between requests — state evaluators cannot see — and compacts long conversations. One condition plays the games as a nearly amnesiac loop, re-deriving everything each turn; the other plays as an agent with working memory. With memory on, effort barely matters (every Provider Adapter row sits between 96.7% and 99.9%); under the Standard harness, effort is the whole game (17.5% to 62.7%).
OpenAI's own launch blog reports the ARC-AGI-3 row as 99.9% — ARC's Provider Adapter result at high effort — against GPT-5.6 Sol at 7.8% and Claude Opus 5 at 30.2%. It carries a footnote: Astra ran on OpenAI's responses-API harness, "which changes two settings to better match real-world performance," settings that "do not specifically target ARC-AGI-3." Our launch coverage gave that footnote a sentence; it deserves more, because OpenAI has been making this argument since July.
OpenAI made this case before Astra existed
On July 29, OpenAI researchers published "How enabling two settings tripled our scores on the ARC-AGI-3 benchmark" — about GPT-5.6 Sol, weeks before Astra shipped. Sol scored 7.8% on ARC's official harness; GPT-5.5 could barely play, at 0.4%. OpenAI's diagnosis was that the harness, not the model, was failing: ARC's setup discarded private reasoning after every action, forcing Sol to "figure out the game anew" each turn, and dropped older observations past roughly 175,000 characters. Enabling retained reasoning and compaction — both standard in ChatGPT and Codex — tripled Sol's public-set score to 38.3% with six times fewer output tokens. "Benchmarks rarely measure AI models in isolation," the post concluded. "They also measure a bundle of less visible choices about API settings, harness design, and prompting."
ARC Prize's response is visible in the September table: it adopted the Provider Adapter as a second, labeled condition rather than disputing the numbers. But ARC keeps the two questions separate: a future AGI, it argues, should solve ARC-AGI-3 under the minimal Standard interface, which stays the apples-to-apples comparison across providers; the Provider Adapter answers how well a model performs when it can lean on the context-management machinery its maker designed. The uncomfortable finding: those answers diverge by 37 points, and on this benchmark a vendor's production stack now outweighs a year of model progress.
The screenshot versus the fine print
The public fight began with a screenshot. A graphic at the top of r/singularity showed Astra at 98.6% against Sol's 7.8% and Opus 5's 30.2% — an insurmountable-looking lead, echoed in headlines like The New Stack's "GPT-6 Astra's score of 98.6% looked like AGI. Then researchers read the fine print." The pushback thread laid out the case: Astra's figure came from a custom harness while the rivals' numbers were Standard-harness runs; effort was unmatched, with Astra at max reasoning against Opus 5 at high; and if custom harnesses count, older models had already posted 99% to 100%. The poster's preferred comparison — both models, Standard harness, matched effort — gives Astra 54.8% against Opus 5's 30.2%: "still a noticeable lift but much less misleading than what OpenAI chose to report." That thread drew roughly 650 points and more than 130 comments.
The counterargument, made best in an r/OpenAI thread, is that "harness cheating" is the wrong frame. Real agent workflows are stateful; if production performance depends on state persistence, compaction, and scaffolding, the model-plus-harness system is what matters. The New Stack's Amanda Caswell lands the synthesis: performance "more often depends on the system around the model." The outlet even added an editor's note — 100% on ARC-AGI-3 means performing at or above the median human baseline, not that the benchmark was "aced." Nobody in the debate is wrong about the mechanics; they disagree about what the benchmark is for.
What saturation does and does not prove
ARC Prize saw this coming. When the benchmark launched in March — we covered the release — frontier models scored below 1% while humans solved everything, and the launch paper stated upfront that saturating it would not constitute proof of AGI. The September post repeats it in ARC's own words: Astra is "a noticeable step-function change in frontier model capabilities," but "we are not claiming that it is AGI." The milestone ARC claims is efficiency, not just completion: with the Provider Adapter, Astra used fewer actions than the median human on 96% of levels — 51.7% fewer per level on average — a result the designers expected to stay out of reach. The limits are stated in the same post: ARC-AGI-3's environments are deterministic, closed-ended, and tightly bounded. The constructive outcome is procedural: the leaderboard will now carry both harness conditions, labeled, with the testing repository and policy public — an admission that benchmark integrity depends on reporting the scaffold alongside the score.
A run now costs more than the argument
The other number in the grid is the bill. A single Astra evaluation ran from $17,000 to almost $50,000 depending on configuration, and ARC notes the inversion that makes the table read strangely: under the Standard harness, more reasoning effort lowered cost — solved games take fewer actions, while failed low-effort attempts burn tokens without progress. Against that, ARC's human baseline testing paid $115 per 90-minute session plus $5 per game — about $12.78 per attempted game, or 0.067 cents if you price the brain's energy as electricity. When one benchmark run costs what a startup's monthly GPU bill used to, replication stops being a hobby. The rival numbers in OpenAI's comparison chart were measured at different times, efforts, and harnesses — which is why the demand for a single neutral standard-harness leaderboard is the most actionable idea this episode produced.
What would settle this
ARC Prize has already committed to the first step: dual-harness, clearly labeled results on its public leaderboard, so no vendor figure floats free of its scaffold again. The rest is replication — a third party running Astra and its rivals on the Standard harness at matched reasoning effort, publishing cost per percentage point alongside raw scores. Until that exists, treat "Astra saturates ARC-AGI-3" as ARC's designers do: a statement about a model inside a very capable harness — a real achievement, and not yet a claim about general intelligence. The 37-point gap is not a bug in the reporting. It is the finding.


