The spoken task is not a fixed script
Qwen’s official repository shows Qwen-Omni observing a scene, proposing a manipulation task aloud and judging the robot’s execution. The team says the robot completes tasks on the fly without a predefined task list. The clips make the idea easy to grasp: the model has to connect an instruction with what is visible, then produce physical actions.
The underlying Qwen-RobotManip system pairs a Qwen3.5-4B vision-language model with an action expert for continuous control. Qwen says the project uses a shared 80-dimensional representation to align robots with different bodies, cameras and action spaces.
From human video to robot data
Qwen reports a training corpus of roughly 38,100 hours, including 24,808 hours of synthetic robot demonstrations generated from 1,933 hours of egocentric human video across 15 robot embodiments. The method retargets human hand activity, removes and inpaints the hand, renders the robot and composites depth-guided imagery.
That scale matters because robots do not all share the same joints, cameras or coordinate systems. Qwen’s alignment framework tries to put different platforms into a common action representation, then uses embodiment prompts and recent observations to adapt a policy without updating its parameters.
What viewers should know
The repository reports strong benchmark results and a real-robot demonstration set, including reactive retries after a slip or failed grasp. Those are Qwen-reported research outcomes; a short edited demo does not establish reliability in homes or workplaces.
The practical detail is also easy to miss: Qwen says there is currently no plan to release the model weights. The public value today is the report, repository resources and demos, which let builders examine how the approach handles unseen instructions, scenes and robot bodies.
Source published June 16, 2026; repository materials accessed October 7, 2026. Coverage is based on the maker’s announcement and demonstration.
