All News
gpt-6-astraroboticsbenchmarkrobocurveagents

GPT-6 Astra scored 95% on a real robot arm — then tied Fable 5.1 when precision mattered

RobotCurve's real-robot eval: GPT-6 Astra scores 95% on a block-into-bowl task versus Claude Fable 5.1's 40% — then both stall at 10% on precision insertion.

Vlad MakarovVlad Makarovreviewed and published
5 min read
GPT-6 Astra scored 95% on a real robot arm — then tied Fable 5.1 when precision mattered

An independent robotics benchmark reports that OpenAI's GPT-6 Astra can drive a real robot arm through an easy pick-and-place task almost every time — 19 of 20 trials versus 8 of 20 for Claude Fable 5.1 — until the manipulation requires precision. RobotCurve, a self-described Public Benefit Corporation measuring frontier robotics capabilities for the public, posted the results September 4 as a follow-up to its earlier Fable 5 and Fable 5.1 robot runs. Anthropic's flagship against OpenAI's, tested by a third party rather than either vendor — and all 120 trials are public, transcripts and videos included. The 95% figure is circulating as a possible sign of AGI; the fuller table is the antidote.

What the robots were asked to do

RobotCurve gave Astra the same bimanual I2RT YAM arms — 6-DoF per arm with parallel-jaw grippers — from the Fable report, under the open-source Inspect Robots harness (MIT) with the same agent policy. Control is coarse: each turn the model sees three camera views (top, left wrist, right wrist) plus proprioceptive state, and replies with an absolute end-effector pose per arm: x, y, z, yaw, pitch, roll and gripper. RobotCurve's Jay Chooi spelled out the interface on X: "We let Astra and Fable control the arms via specifying the end-effector pose and passing it through an automatic IK solver." The LLM picks poses; the solver does the motion. The two task prompts, verbatim:

  • "Pick up the red block from the table and place it inside the bowl."
  • "Pick up the round blue puzzle piece by the knob at its center and place it into the matching circular groove in the board."

Each model ran 20 trials under an agent policy: medium thinking effort, a 20-LLM-call budget, 25% arm speed, default guardrails. A human operator graded every run on the highest stage it reached.

The scorecard

Runner-measured numbers from an independent organization, not a vendor table — though the runner also graded with the model known (more below). Verbatim from the eval page:

TaskModelMean stageCompletionsRateOutput tokens/runEst. cost/runMinutes/run
Block into bowlFable 51.301 / 205%19.2k$2.698.2
Block into bowlFable 5.12.408 / 2040%12.9k$2.126.8
Block into bowlGPT-6 Astra3.9519 / 2095%2.1k$0.942.5
Puzzle into grooveFable 51.500 / 200%16.3k$2.637.9
Puzzle into grooveFable 5.12.352 / 2010%10.5k$2.185.9
Puzzle into grooveGPT-6 Astra2.002 / 2010%2.7k$1.363.4

Costs apply list price ($10 / $50 per million input/output tokens for all three models) to wire-level tokens, not billed ones; OpenAI auto-cached about a fifth of Astra's input without discounting it, per RobotCurve, so Astra's cost is if anything overstated.

The 95% is one task; the wall is another

On the bowl task Astra is not just more reliable than Fable 5.1, it is cheaper per outcome: about six times fewer output tokens (2.1k versus 12.9k), roughly 2.3 times lower estimated cost ($0.94 versus $2.12) and 2.5 minutes per trial versus 6.8. On the puzzle task it keeps those advantages (2.7k versus 10.5k tokens, $1.36 versus $2.18) and ties Fable 5.1 at 2 of 20 completions. RobotCurve's own description of Astra's failure: "It reaches the groove and stalls at the same final step Fable does." Its mean stage there, 2.00, trails Fable 5.1's 2.35: Fable 5.1 more often reached the positioned stage before failing to seat the piece. Insertion needs contact, force and millimeter tolerances — and every model hit the same wall, the least AGI-like result in the report.

The confessions are in the report

RobotCurve lists its methodological weaknesses rather than burying them:

  • Astra's trials ran two days after the Fable trials, not interleaved with them.
  • The bowl comparison used different rigs (Astra on rig-1, the Fable models on rig-3) because rig-3 was unavailable; the puzzle task ran on rig-4 for all models.
  • Grading was operator-judged with the model known, so scores are open to unconscious bias.
  • Objects were reset by hand between trials.

None of this explains away a 19-of-20 versus 8-of-20 gap, but it is the kind of condition that recently moved Astra's scores on a static benchmark: ARC's harness re-run of ARC-AGI-3 swung the model from 17.5% to 99.9% with harness and effort settings. Physical trials add rig and operator variables on top.

Faster, cheaper, slower than the hype

The "real-time robot control" framing needs three corrections. Arms ran at 25% of their speed. The videos' timers strip thinking time: the page labels its best-run clips as showing "real elapsed time with thinking pauses removed," and the Reddit demo's timers claimed 20.7x the same way. And the end-of-year claim is extrapolation, not result: "If trends hold, LLMs could control robot arms in real time as soon as the end of this year, or by 2029 at the latest," Chooi wrote in his announcement thread, which led with "GPT-6 Astra scored 95% on a robot control task, up from Fable 5.1's 40%, with 6.2x fewer output tokens at 2.3x lower cost" and drew roughly 5,200 likes. r/singularity's "Signs of AGI?" thread drew roughly 535 points and more than 120 comments — engagement ahead of what the second task shows. For the launch-week Astra story, the report adds what vendor tables cannot: an independent lab measuring the shipping model at medium effort, more reliable and far cheaper than its rival at coarse manipulation.

What would settle it

Interleaved trials on one rig with blinded grading would remove the largest confounds. Harder tasks would test the wall: tighter insertion tolerances, deformable objects, force-sensitive contact instead of pose-by-pose control. Real-time control means closing the loop without a 25% speed cap or timers that remove thinking pauses. Until then, the eval page carries the honest summary: Astra puts a block in a bowl nearly every time, and stalls at the same groove as Claude when the task asks for a steady hand.

Related Articles

Scroll down

to load the next article