An agent that can inspect the robot it controls
The researchers connected GPT-6 Astra to an XLeRobot through a software interface and studied an elevator-button task. They supplied different kinds of body knowledge and records of previous attempts, then measured how those inputs changed the agent’s work in simulation and on hardware.
The main idea is that the model can interpret machine-readable robot geometry, read synchronized image/action records, and adapt movements from current camera views. It is not a dedicated learned controller replacing the agent; it is an experiment in giving a general multimodal coding agent structured ways to understand its body and reuse experience.
The measured result
In 30 fixed-start simulation trials, complete geometry and camera information reduced mean completion time 57.4% against the baseline. Reusing experience also helped at starting positions displaced 10 to 100 centimeters in the tested setup.
On the physical robot, all 12 trials reached operator-confirmed gripper contact with the button. At a shared nominal start, simulation geometry reduced mean time by 53% and simulation experience by about 50% compared with the no-resource baseline. The success condition was operator-confirmed contact, not a verified button press.
What carries over—and what does not
The experiment suggests that an agent can reuse useful pieces of a prior attempt rather than replaying a whole movement sequence. It can take a known pose as a starting point, then use current images to adjust its approach.
This is a small, carefully instrumented research task, not proof that Astra can operate a general-purpose robot across arbitrary buildings. The paper is unusually useful because it releases prompts, trial-level data and skill implementations, making it possible to inspect what “robot agent” meant in the experiment.
Source published September 24, 2026. Coverage is based on the maker’s announcement and demonstration.
