A human video becomes a robot’s missing experience

XPENG Robotics’ AnyWorld starts with an egocentric human interaction and recomposes it into a robot-domain rollout. Its controls separate the action from camera movement and target embodiment, so the same interaction can be rendered with a different robot, viewpoint or scene. The project says it does not need paired human-and-robot clips of the same action.

That matters because robot datasets often contain a narrow set of bodies, objects and task states. AnyWorld is designed to generate more varied robot-facing examples, including targeted examples for a behavior a policy has not learned.

The banana already in the box fooled the baseline

XPENG’s clearest example begins with a robot-policy shortcut: when a banana is already visible inside the basket, the baseline stays still—even though the instruction requires placing another banana there. AnyWorld uses a human interaction as the motion prior, then re-embodies and re-scenes that interaction into the missing robot state.

The accompanying IRON clip shows the robot pick up the remaining banana and place it in the basket. The point is not simply that a video model can make a plausible robot scene; the generated example is used as targeted supervision for a physical robot policy.

A promising result with a small trial count

On the project page, XPENG reports an IRON real-robot grasp evaluation of 20 trials: task success rises from 20.0% at baseline to 55.0% with AnyWorld data. The same page reports a smaller improvement in a RoboCasa GR1 pick-and-place evaluation, from 49.8% to 54.6% across 18 tasks.

Those are author-reported research results, not a large independent benchmark. The real-robot sample is only 20 trials, and the result does not establish how broadly the generated cases transfer. Still, the example is unusually concrete: the method targets a visible failure state, generates robot-domain experience for it, then tests whether the robot resumes the task.

Why this is different from just making more robot video

AnyWorld separates the interaction from its realization. It can preserve the motion while changing the body, camera path or surrounding objects, and its policy examples can add rare states such as a partially completed task. This makes it a data-generation approach aimed at policy gaps, as well as a world-model research project.

The research is from XPENG Robotics, whose team has also published its model code. The project page and paper provide the architecture, evaluation setup and example clips for readers who want to inspect the method directly.

Explore the original source ↗

Source published September 1, 2026. Coverage is based on the maker’s announcement and demonstration.