A cooking clip becomes a robot trajectory

The Do as I Do research team takes ordinary single-camera footage of a person handling an object, estimates the hand-object interaction, then retargets that motion to a multi-fingered robot hand. The project video shows a human cooking beside a robot reproducing a stirring motion.

The point is not that the robot copies pixels. The system reconstructs the object and hand interaction, then optimizes a physically executable trajectory for the robot embodiment.

Why the conversion is difficult

A human video does not directly specify the forces, finger positions or timing a robot needs. Even a small error in hand-pose estimation can make a grasp impossible to execute.

The paper addresses that gap with two stages: reconstruct the hand-object interaction from monocular RGB footage, then retarget the result onto a robot hand. The authors demonstrate motions with different objects and grasps.

What is demonstrated

The team’s project page shows ten real-world motions, reconstruction overlays and retargeting in MuJoCo. Its paper reports better hand-object interaction estimation and dexterous trajectories than earlier approaches on the authors’ evaluation datasets.

This is a research pipeline for robot-learning data, not an off-the-shelf robot skill downloaded from any video. Each motion still depends on the reconstruction and retargeting stages working for the objects, camera views and robot hand involved.

Why it could matter

Robot manipulation data is expensive because people often have to teleoperate a robot or build a specialized capture setup. If ordinary videos can be converted into physically valid examples, the web becomes a much larger source of demonstrations.

The big question is whether those reconstructed trajectories remain reliable across many objects and real robot hands. The project makes that translation visible: human footage on one side, a robot hand reproducing the motion on the other.

Explore the original source ↗

Source published 2026-06-17. Coverage is based on the maker’s announcement and demonstration.