All News
gpt-6-astraopenaigamingagentscomputer-use

GPT-6 Astra beat Portal — after $571 in API calls

A user's GPT-6 Astra agent just autonomously finished Portal in under 24 hours for $571.18 in API calls. We break down the run, and why skeptics object.

Vlad MakarovVlad Makarovreviewed and published
7 min read
Mentioned models
GPT-6 Astra beat Portal — after $571 in API calls

GPT-6 Astra has beaten Portal — every chamber of it, down to the last confrontation with GLaDOS. The 2007 Valve game, in which the player-character Chell escapes a murderous research facility by shooting linked holes in the walls, is one of the most analyzed titles ever made, and Portal needed no help being beaten; it has had tool-assisted speedruns for years. What happened on September 5 is different: the user cozyblaze announced on X that OpenAI's GPT-6 Astra finished the full game autonomously, a claimed first for a general-purpose frontier model. The run took roughly 24 hours, 3,336 tool calls, and $571.18 in API fees. Whether that is a milestone or a stunt depends on how much weight you give to the word "autonomously."

A 24-hour run, announced like a record

cozyblaze's announcement read like a cross between a speedrun claim and a lab report. The post, published at 11:40 PM on September 5, tied the run back to 2016, when one of OpenAI's stated technical goals was to solve a wide variety of games using a single agent — and presented Portal as that ambition, finally cashed in. The coverage followed within a day. VideoCardz tallied the 3,336 tool calls and the $571.18 bill, and The Verge's write-up framed it as a bargain, noting the model beat "the full game in under 24 hours for less than $600." The r/singularity thread "GPT-6 Astra Has Beaten Portal, Becoming the First Model to Achieve This" drew roughly 2,300 points and more than 220 comments in its first day, most of them arguing about the asterisks.

Pause, screenshot, decide, repeat

The mechanics matter more than the milestone. The model never sat in front of a monitor with a controller; it worked through a connected toolset built around MCP and a modified SourcePauseTool (SPT), which fed it screenshots of the game and executed the input sequences it wrote. The game stayed paused while the model thought; each decision was a full cycle of frame, decision, and brief execution. cozyblaze's own condensed recording of the run, posted to YouTube on September 6 as "GPT-6 Astra Plays Portal," is explicit about the arrangement: "The game stays paused while the model thinks; once the model sends an input sequence, SPT unpauses and executes it." The video had roughly 49,000 views and 1,100 likes within a day.

The model saw the game through default 360p screenshots plus a position observation covering player facing and roll, and the setup supplied its coordinates directly. The harness also leaned on the context-management system Astra introduced at launch last week to keep its working state intact across the marathon. The ledger from cozyblaze's own stats:

  • input tokens: 430.5M, including cached reads
  • output tokens: 1.6M, 432.1M total
  • context load hovered near 50%: roughly 138.0k of a 258.4k window
  • tool calls: 3,336
  • wall time: about 24 hours, including reasoning pauses
  • environment: Portal build 5135, Source Unpack

When the clock stops, is it still playing?

Reddit's skeptics did not wait for the video. "It freezes the game at thinking time. That's not really beating the game," read one of the thread's sharper comments, and the objection is hard to wave away: with the game paused for every decision, Portal becomes a turn-based exercise in screenshot interpretation — closer to a visual reasoning loop than to the real-time play the headlines imply. The second line of attack was contamination. Portal is nearly two decades old and among the most documented games ever shipped, so, as a commenter put it, "Portal may be in the training data," meaning the model could be reconstructing a memorized route through the chambers rather than discovering one. A related point drew the most technical heat: each tool call is code controlling a TAS tool, with the agent writing whole input sequences for a modified Source engine to execute — and tool-assisted runs have been finishing Portal for years on a fraction of this compute. Comparisons to DeepMind's game-playing systems followed naturally. Not all of it was dismissive: one commenter, Zermelane, measured Chell's velocity in the footage to reverse-engineer the portal-firing angles the model chose.

A $571 playthrough is not a benchmark

Then there is the price. A competent human finishes Portal in an afternoon; this run cost $571.18 in API fees, and the token ledger explains why. With 430.5 million input tokens, the model was repeatedly re-reading its own history, the context window sitting near half its 258.4k capacity at any moment. VideoCardz's headline captured the tension — "AI model successfully completes Portal but it costs $571 in tokens" — and cozyblaze conceded the point in a follow-up reply: "There are still many problems to solve, and it's not a proper benchmark per se." The energy angle sharpens it further: every token cost real electricity and, through the cooling systems behind the inference, real water — The Verge's aside that the feat "probably only required a small swimming pool's worth of water to boot" was aimed exactly there. A TAS input file beats the game with near-zero marginal compute, which is why skeptics argue this run says more about the harness, the pause, and the budget than about a new game-playing intelligence. It is also why the milestone survives only in a qualified form: a general-purpose model reaching the end credits by reading screenshots and issuing commands, without a human hand in the loop, is new. Whether that counts as playing is a separate question.

The harness is the story, again

Read generously, this is the 2016 goal cozyblaze invoked: a single agent, trained on language rather than on Portal, took in pixels and produced actions that finished a game it was never programmed for. Read skeptically, it is the same lesson as the ARC-AGI-3 dispute that opened the week, covered here: the harness is part of the score. Astra's ARC results swung by dozens of points depending on whether the harness preserved its reasoning state between requests, and the Portal setup — pause-and-resume thinking, a coordinates feed, a fresh context-management system — does comparable work. The model-and-harness combination may be the capability that matters, and this run is a demonstration of that combination, not of the model alone. For OpenAI, whose launch materials already leaned on computer-use demos, the run lands as free marketing for agents that operate real software: the flagship shipped days earlier, and here it is filing its own speedrun.

The worst model that will ever do this

cozyblaze's own summary doubles as the best rebuttal to the hype: "This is the worst model we'll ever get." Every future flagship should repeat the run cheaper, faster, and eventually without the pause at all. That is the trend line worth watching — not whether a frontier model can finish a 2007 game with an assistive harness, but how quickly the assistance disappears: no supplied coordinates, no frozen clock, no $571 bill, and finally a game written after the training cutoff, which would settle the contamination argument outright. It has now been beaten by a model that reasoned its way through, slowly and expensively, one paused screenshot at a time. Fans will argue about the asterisk for years; the argument about the direction of travel is already settled.

Related Articles

Scroll down

to load the next article