From first-person video to a robot hand
XPENG Robotics introduced IronMind as a humanoid manipulation system pretrained on more than 10,000 hours of curated egocentric human video. The project keeps demonstrations in the camera frame instead of trying to reconstruct a human torso frame, then transfers the learned policy to its IRON-R01 humanoid.
The practical test is whether the robot can use its hands on objects and instructions it did not see during training. The project page labels its example clips fully autonomous and says they are shown at 1× speed. The clips are demonstrations from the research team, not independent replications.
Six tasks make the claim easier to judge
The examples include putting a toy car in a basket, picking up a teapot, moving an orange to a basket, handing a basket to a person and following a left-basket instruction for grapes. A blue-cup handover adds a color-and-action prompt. The tasks make it possible to inspect grasping, target selection and handoff separately.
XPENG reports that its 10,000-hour model scored above 40% average success across six real-robot out-of-distribution tasks, with individual task results from 40% to 90%; models trained on up to 5,000 hours stayed below 12% on average. Those are the company’s experiment results, not a general-purpose reliability guarantee.
Why the data scale is the story
The visible clips matter, but the bigger result is the bridge from everyday human hand video to robot control. Human footage is abundant; robot demonstrations are expensive. IronMind’s approach tries to make that gap smaller without requiring a matching robot-body trajectory for every human action.
For builders, the source page includes the task clips, training details and paper. Watch the sequence itself: the robot selects an object, reaches, grasps and places or hands it over. That is a clearer test of dexterous behavior than a static model score.
Source published September 30, 2026. Coverage is based on the maker’s announcement and demonstration.
