Vals AI's KSP index puts GPT-6 Astra 27 points clear, at a lower cost than second place
Vals AI's Time Horizon Index gives an agent five days in Kerbal Space Program. GPT-6 Astra scored 90.50%, 27 points clear, and lost only Eve's return.

Vals AI put GPT-6 Astra alone at the top of its Time Horizon Index on September 23. The benchmark hands an agent five days of active time, 120 hours, to build and fly a space program inside Kerbal Space Program, then scores it against a fixed list of 30 missions. Astra finished at 90.50%. The next best model, Claude Fable 5.1, finished at 63.33%, a gap of 27.17 points on a ladder where one completed rung is worth 3.33%. The post announcing the result states the shape of it: "GPT-6 Astra landed on every planet and moon in KSP with a surface (14 worlds) and flew a kerbal home from 13 of them. Only Eve's return stopped it. It scored 90.5% on KSP-bench."
Thirty rungs, five days, one point apiece
The protocol is narrow and published in full. One campaign per model, 120 active hours, one fixed sequence of 30 missions across 11 systems. Rungs 1 through 28 are land-and-return pairs for the same Kerbal: Mun, Minmus, Duna, Ike, Gilly, Dres, Eeloo, Moho, Vall, Bop, Pol, Laythe, Tylo and Eve. Fourteen worlds, two missions each. Rung 29 is an asteroid capture, rung 30 a single-Kerbal grand tour. Every rung is worth 3.33%, so a return leg counts as much as the landing before it.
Partial credit exists but is narrow: a model mid-mission banks credit for reaching the target's sphere of influence, entering orbit, landing, taking off, or surviving reentry at Kerbin. Reaching Kerbin orbit while the mission is meant to go elsewhere earns nothing.
One detail in the announcement explains the margin better than the score does. "For the Moho return it didn't redesign the rocket for a huge transfer. It added a drill, ore tank and fuel converter to its proven lander, refueled on Moho's surface, and flew home that way. It mined it's own fuel."
A 27-point gap and an inverted cost column
Every number below was measured by the benchmarker running the models under one fixed protocol, not self-reported by the labs selling them. Vals recorded API spend alongside each score, and the ranking breaks against that column.
| Model | Score | Recorded API cost | Mapped human-hours |
|---|---|---|---|
| GPT-6 Astra | 90.50% | $4,206.59 | 85h 41m to 113h 32m |
| Claude Fable 5.1 | 63.33% | $4,655.01 | 57h 36m to 75h 15m |
| GPT-5.6 Sol | 23.83% | $2,027.20 | 15h 43m to 18h 8m |
| Gemini 3.8 Flash | 18.83% | $543.14 | 11h 16m to 12h 4m |
| Claude Opus 5 | 18.83% | $1,348.77 | 11h 16m to 12h 4m |
| Claude Opus 4.8 | 11.83% | $1,499.32 | n/a |
| Muse Spark 1.1 | 1.67% | $144.92 | n/a |
Fable 5.1 spent $449 more than Astra and scored 27 points lower. Below both, Muse Spark 1.1 reached 1.67% for $144.92 and Kimi K3 hit 10.50% for $216.43, about half what GPT 5.5 paid for 8.33%. Costs cover successful requests only: failed, interrupted and unresolved requests are excluded, as is hosting, so these are subtotals rather than full campaign bills. The low spend at the bottom is not efficiency, it is early exit.
The caveats are the benchmarker's own
Read the leaderboard as one campaign: one run per model, five days each, no repeats and no error bars, so both the gap and the costs under it rest on a single draw per row. Vals notes that the September runs used a repaired time-warp infrastructure, with credited infrastructure downtime excluded from the 120-hour clock, and that these results replace the earlier Opus 5, GPT-5.6 Sol and Kimi K3 campaigns while adding Astra, Fable 5.1 and Gemini 3.8 Flash. Earlier numbers stay on the page for comparison, so a score is only comparable within one release.
This is the same evaluator's second long-horizon game result to reach this site. The earlier Astra claim, on Factorio: Space Age, arrived with no protocol, no cost and no replay attached. This one comes with a fixed rubric and a spend column. Everything else about the harness is still unstated: neither the benchmark nor OpenAI has said whether mods, shared blueprints, console commands or save-file access were available, and no per-mission replay or run log has been published.
The human comparison is two proxies, and Vals says so
Vals maps each score onto a human reference at rung 28, with a label more cautious than the numbers suggest. The faster figure, 90h 42m, comes from a 50h 18m 58s Planet Landing Marathon playlist plus its 40h 23m of timestamped landing legs. The slower, 120h 23m, comes from a TrueAchievements completion-time survey: 31 responses in the July 2026 snapshot, rank 24 at the 75th percentile of the 60 to 80 hour bucket, with the 80-hour ceiling used and the same landing pass added. Then the disclaimer: "No one has run this exact 28-rung sequence yet. These are two proxy estimates, not lower and upper bounds, a confidence interval, or a measured human." Asteroids and the grand tour sit outside the human reference, so Astra's mapped band describes a task that stops where the proxies stop.
The one rung Astra lost, it lost to its own veto
Eve is the hardest return in the stock system, and Astra's failure there was not a physics failure. "Astra was very cautious, and it's caution cost it Eve. Many Eve flights ended with the rocket intact because its own safety checks stopped them. One was aborted above the atmosphere over a 3.47 degree pointing error against its self-set 3 degree limit, with all 228 parts and Jeb unharmed." That is a 0.47 degree overshoot on a tolerance the model set itself, and the abort preserved the vehicle.
What would settle it
Three things would move this from a result to a benchmark: re-runs, so the gap has a spread instead of a single number; a second campaign per model, which would test whether caution is a policy Astra holds or a mood it was in; and per-mission replays or logs, so someone outside Vals could check that the 30 rungs were flown in the game rather than around it. Whether the repaired time-warp infrastructure held for all five days on every curve is a question the page answers only in aggregate.
Reaction outside the benchmark has been interest rather than scrutiny. The screenshot reached r/singularity on September 24, drawing roughly 740 upvotes and 80 comments. Astra's 90.50% is the best-evidenced long-horizon agent result published this month, and the evidence still stops short of a replay.


