All News
roboticsgpt-6-astrahumanoidstanfordresearch

GPT-6 Astra cleaned a kitchen it had never seen. The demo is real; the autonomy headline is doing a lot of work.

Stanford and Caltech let GPT-6 Astra run a Unitree G1 in an unseen kitchen. We separate the impressive demo from the autonomy claim behind its viral headline.

Vlad MakarovVlad Makarovreviewed and published
7 min read
GPT-6 Astra cleaned a kitchen it had never seen. The demo is real; the autonomy headline is doing a lot of work.

On September 26 a video appeared on r/singularity claiming that GPT-6 Astra can now control a humanoid robot in a room it has never seen, remember where objects are, clean up across the room and fetch things later from vague human requests. The post drew roughly 1,400 points and more than 250 comments over the weekend. The research behind it is HomeBody, a Stanford and Caltech project page published this month, and the page is both more interesting and more carefully scoped than the post that carried it. It is also a demo writeup rather than a measurement, and the difference matters more than the headline.

What the demo actually shows

HomeBody comes from the Movement Lab at Stanford, by Gio Huh at Caltech with Cayden Gu, Takara E. Truong, C. Karen Liu and Guy Tevet, the last two listed as equal advisers. The setting is a kitchen the robot had never entered, at the Stanford Robotics Center, and the robot is a Unitree G1 guided by GPT Astra. The project page quotes its two task instructions verbatim.

"Clean up all of the coffee bags and put them in the middle, and throw away all of the milk and orange juice cartons that have gone bad."

"I forgot my medicine, can you get it for me? Also throw out the bad carton while you are at it."

The second is the harder one. The medicine starts out of view, so the robot has to locate the drawer from stored keyframes, open it with the right hand, hand the bottle over, then switch arms and discard the carton with the left. The Decoder covered the project on September 27 under a headline about plugging Astra directly into a robot.

A hybrid stack, not a brain transplant

The page sets up the standard picture first: a System 2 VLM reasons, a learned System 1 vision-language-action model turns that into commands, and a System 0 controller executes motion. HomeBody asks whether the learned VLA is still needed, and instead hands a replaceable frontier VLM a library of reusable motor skills — pick, place, open drawer, pick from drawer, navigate — with observations and execution outcomes returned to the model. Coverage has read this as the model controlling the robot. The page describes something narrower.

LayerWho does it
DecidesGPT Astra: which skill, which target, which hand, what next
MovesPretrained AMO lower-body policy at 50 Hz, arm and hand commands at 250 Hz
SeesD435i stereo depth, SAM 2.1 tracking, Super Odometry localization

The VLM's output is a structured tool call, not a motor command. For a pick it names a point in the ego image normalized to 0-1000 and which hand to use; segmentation, stereo depth, analytic grasp prediction, spline-based arm planning and inverse kinematics do the rest, and navigation follows a planned route in the reconstructed map. The same division of labor ran in DrivingBench, where Astra picked poses for a real car and a solver turned them into motion. Target selection by a frontier model is a real capability claim. Body control is a different one.

The setup bill comes first

Before any instruction, the robot is told to explore: "You are a kitchen robot, please explore the space!" It wanders the room collecting 0.5x iPhone video, D435i observations, LiDAR scans with SLAM, joint poses and waypoints chosen by Astra. Then Astra is used a second time, as the Real2Sim agent, to build a digital twin of that kitchen in NVIDIA Isaac Sim out of the robot's own data, complete with a reconstruction you can click into and compare against LiDAR. Only then does a plain-language instruction get planned as skill calls.

The "no environment-specific training data or additional policy learning" line is true of this kitchen, and it does not mean the system arrives with nothing. The skill library is pretrained and engineered by its authors, and the room-specific work is displaced into the exploration pass plus the digital twin, which the authors themselves bill as setup time and API costs. Free autonomy is not what the page claims, and it is not what the videos show.

What the page does not report

There are no trial counts anywhere on it. Two tasks, one kitchen, one embodiment, and no repetitions, success rates or failure breakdowns — no statement of how many attempts a task took, or how often a person had to step in. The skill previews note that drawer opening and grasping "include recorded reattempts" and that the drawer-pick clip "starts with the drawer open and keeps the retries continuous," which tells you retries are real without saying how many there were. The runs that exist are, by construction, the ones worth publishing.

That gap is exactly what third-party robotics evaluators were built to close. RobotCurve ran 20 trials per model per task on real arms — 120 in total — graded every run and published each transcript, camera recording and trajectory, and it reported that Astra tied Claude Fable 5.1 at 2 of 20 on the precision task rather than burying it. A project page showing one kitchen going well is a different kind of document.

The press framing has a second soft spot. The code repository that the-decoder links says, plainly, "Code coming soon." What is public is a README, a docs folder and six commits.

What is actually new here

Two elements are more than incremental. The first is persistent spatial memory grounded in the robot's own sensor data: keyframes stored in a single coordinate frame shared by the SLAM map and the reconstructed twin, so an object that has left the field of view can still be found and an underspecified request resolved. The second is the skill interface itself, where execution results flow back to the VLM so it can re-target or change plan, and where a new skill — a learned policy, a classical algorithm, another controller — plugs in behind the same contract. That architecture is the part most likely to outlive the demo.

The limits, in the authors' own words

The scope section is unusually direct: "Real2Sim reconstruction adds setup time and API costs. Task length is also constrained by the humanoid's reach, manipulation capabilities, and hardware endurance, including finger-servo overheating during extended operation. GPT Astra's reasoning latency introduces pauses between skills." It adds that the local stack requires an RTX 4090 laptop GPU, with heavier perception or more skills likely needing more compute. That is a list of pressures that get worse as tasks get longer, not better.

What would settle it

Three things, all cheap for a lab that already has the kitchen. Repeat the two tasks many times and publish the counts, across several unseen rooms rather than one. Test the same pipeline with a different frontier VLM, because the architecture panel labels Astra as "the recorded model," which implies the swap is a design property that nothing on the page demonstrates. And get an outside group to replicate it on hardware they built, since a kitchen in the lab that built the system is not an unseen environment in any adversarial sense.

Then the claim in that forum title would be a result. For now it is a well-made demo wrapped around a genuinely interesting architecture, and the question it asks about what the model layer still needs to be told is the right one.

Related Articles

Scroll down

to load the next article