Researchers Gio Huh, Cayden Gu, Takara E. Truong, C. Karen Liu, Guy Tevet from Stanford University and Caltech (The California Institute of Technology) have introduced HomeBody, a new robotic system that enables humanoid robots to autonomously explore unseen spaces, maintain persistent spatial memory, and execute complex household tasks without task-specific training.
In the report, they share that HomeBody bypasses the traditional learned Vision-Language-Action (VLA) pipeline. Instead, a frontier Vision-Language Model (GPT Astra) directly plans and triggers a modular library of motor skills; such as walking, opening drawers, and picking up items. After a brief initial exploration pass, the robot constructs a geometrically and semantically accurate 3D digital twin in Isaac Sim (a simulation platform built on NVIDIA Omniverse) to reason about objects outside its immediate line of sight.
Tested on a Unitree G1 humanoid in a brand-new kitchen, HomeBody successfully organized scattered items and retrieved hidden medicine based on vague verbal requests.
This modular approach significantly lowers the barrier for deploying autonomous general-purpose household assistants.