On September 27, 2026, researchers from Stanford and Caltech published HomeBody, a system that does something the robotics field has mostly avoided: it wires a frontier vision-language model directly to a humanoid robot, with no separately trained control policy sitting in between. The model, GPT Astra, reasons about a room and calls a small library of hard-coded skills to act. In the demonstrations, a Unitree G1 humanoid explores an unfamiliar kitchen, builds an internal map of it, and then tidies up or fetches an item on command.
HomeBody is a research project, not a product, and the honest framing is that it is interesting for its architecture rather than for solving home robotics. But the architecture is exactly the part worth understanding, because it stakes out one side of the central debate in embodied AI: whether the path to capable robots runs through massive end-to-end learned control, or through frontier reasoning models orchestrating simpler pieces.
The design in one sentence
HomeBody uses a general reasoning model as the brain and a fixed set of five skills as the hands. Instead of training a neural network to map camera pixels straight to motor commands, the researchers let GPT Astra plan in language and call composable skills, pick, place, open drawer, pick from drawer, and navigate, to carry the plan out.
A frontier model as the brain, a fixed skill set as the hands
HomeBody puts a general reasoning model directly in charge, with no separately trained control policy in between. The model plans; five hard-coded skills execute.
The robot walks an unseen room, collecting camera video, depth, LiDAR and SLAM data, and its own joint poses.
›From that data the model builds a digital twin of the room in Nvidia Isaac Sim, logging objects and locations as spatial memory.
›Given a plain-language task, the VLM reasons over the twin and calls composable skills, correcting step by step.
Architecture per the Stanford TML HomeBody project page, September 2026.
The pipeline has three stages. First the robot explores an unseen room, gathering camera video, depth, LiDAR and SLAM data along with its own joint poses. Then it uses that data to build a Real2Sim digital twin of the room inside NVIDIA Isaac Sim, logging where objects are as a form of spatial memory. Finally, given a plain-language instruction, the model reasons over that twin and issues skill calls, correcting itself step by step as it goes. The two demonstrated tasks, "tidy the kitchen" and "retrieve the medicine," were chosen to exercise different capabilities: multi-object cleanup with sequencing and navigation in the first, and occluded-object retrieval that leans on the spatial memory in the second.
Why this architecture is a real position, not just a demo
The interesting claim in HomeBody is not that a robot cleaned a kitchen. It is that a general-purpose reasoning model, one never trained to control a robot, can plan physical tasks well enough to drive a humanoid through them in a space it has never seen, given only a library of primitive skills and a map it built itself.
The bet is that a robot's intelligence can be borrowed from a general reasoning model, rather than learned from scratch on robot data.
The architectural claim behind HomeBody
That is a bet against the dominant paradigm. Much of frontier robotics has pursued large end-to-end policies trained on enormous amounts of robot interaction data, on the theory that physical competence has to be learned in the same medium it is used. HomeBody argues the opposite is at least partly true: that a frontier model's reasoning, built from internet-scale data, can transfer to the physical world if you give it clean interfaces to sense and act. The appeal is obvious. Skills are reusable and inspectable, the reasoning improves for free every time the underlying model improves, and you sidestep the need to collect a mountain of robot-specific training data. If it holds up, it is a far cheaper path to general behavior than training a bespoke policy per robot and task.
Where it breaks, and why that matters
A neutral read has to be just as clear about the limits, and the researchers are candid about them. The paper reports no quantitative success rates, which is the single most important caveat: without measured reliability across many trials, the demonstrations show that the approach can work, not how often it does. That gap is exactly where impressive robot videos have historically oversold real capability.

The practical failure points are just as telling. The model's reasoning latency introduces pauses between skills, so the robot thinks visibly between actions rather than moving fluidly. Building the Real2Sim twin adds setup time and API cost before any task can begin. The hardware itself is a limit: the demonstrations ran on a single RTX 4090 laptop GPU, and the robot's finger servos overheated during extended operation, capping how long it can work. Task length is bounded by the humanoid's reach, manipulation range and endurance. None of these is fatal to the idea, but together they explain why this is a research result and not a home appliance. The reasoning may generalize; the body, the latency and the economics do not yet.
How to read it against the rest of embodied AI
HomeBody lands in the middle of an active argument about world models and robot learning. One camp holds that language and vision models are the wrong substrate for physical intelligence and that robots need learned world models grounded in interaction. Another holds that frontier reasoning models are general enough to orchestrate the physical world through the right interfaces. HomeBody is a concrete, inspectable data point for the second view, and its value is partly that the code is public, so others can test how far the approach stretches beyond two curated kitchen tasks.
The broader lesson is one that recurs across AI right now: capability increasingly comes from composing a strong general model with good tools, rather than from training a specialist system end to end. That is true whether the tools are robot skills, software functions or data sources. The same principle shows up wherever people build on AI: a capable general model plus clean, swappable interfaces often beats a monolithic bespoke system, and it keeps the work from being locked to one implementation. It is the reasoning behind model-agnostic software tools such as Metir AI, which treat the frontier model as an interchangeable brain wired to the tools a task needs, the same shape HomeBody uses for a robot, applied to knowledge work.
The takeaway
HomeBody is a small, honest experiment with an outsized idea: that you can borrow a robot's intelligence from a frontier reasoning model instead of training it from scratch, if you give the model a map it builds itself and a handful of reliable skills to call. The demonstrations are genuinely novel and the architecture is a clean statement of one side of embodied AI's central debate. The absence of success metrics, the latency, the overheating servos and the setup cost are equally real, and they are why the right description is a promising research direction rather than a solved problem. What makes it worth watching is that the reasoning half improves automatically as frontier models do, so the same robot could quietly get more capable without new robot training at all.
The pattern behind HomeBody, applied to your work
A strong general model plus clean, swappable tools beats a locked-in specialist system. Metir AI is built on that idea, keeping your work portable across the leading models so the brain can be upgraded without rebuilding everything around it. Try Metir AI free.
Sources:
- HomeBody: A Humanoid That Explores, Remembers, and Acts on Its Own | Stanford TML
- Researchers plug GPT-6 Astra directly into a robot and let it clean up an unfamiliar kitchen | The Decoder
- Stanford's HomeBody wires GPT Astra to a Unitree G1 humanoid | AI Weekly
- Unitree Robotics | Wikipedia
Image credits
Header image: Unitree G1 humanoid robot, CC0, via Wikimedia Commons. In-body image: domestic kitchen interior, via Wikimedia Commons.

NVIDIA