A robot has to move its whole body to finish the task
Google DeepMind’s Gemini Robotics 2 is aimed at tasks that require more than a gripper responding to a command. In one example, an Apptronik Apollo humanoid is asked to put a watering can into a green bin on a lower shelf: it walks to the table, picks up the can, carries it across the room and places it. The payoff is easy to understand because the robot has to coordinate navigation, balance, grasping and placement.
The demonstrations extend to tying knots with a five-fingered SharpaWave hand, packing objects with a Franka platform, and having different robots work together. These are capability demonstrations published by Google DeepMind, not independent measures of reliability or speed. The source page links the individual videos for each task.
Three models divide the work
Gemini Robotics ER 2 acts as the embodied reasoning model: it interprets a request, understands the scene, plans a multi-step task and coordinates robot actions. Gemini Robotics On-Device 2 is a vision-language-action model designed to run locally on robot hardware. The system also includes a model for direct control of robot movements.
That split resembles a planner paired with a controller. The reasoning system decides what should happen next; the action model translates the plan into movement suited to a particular robot body. The architecture is intended to make a robot more adaptable than a fixed script, while still requiring the controller to respect the robot’s sensors, limits and safety rules.
A few hours of adaptation is progress, not instant transfer
Google says the on-device model can be adapted to a new bi-arm robot embodiment in a few hours, typically with fewer than 200 examples. It shows work on Dexmate, SO101 and Trossen platforms with different shapes, sensors and degrees of freedom. If that holds in practice, teams may spend less time rebuilding a policy from scratch whenever hardware changes.
The examples remain bounded. A short adaptation run is not the same as a guarantee that one robot can safely perform arbitrary tasks in an unfamiliar home or factory. Google itself says the robots need to improve movement speed, and the models are still being offered through a mix of AI Studio, private preview and early-access programs.
What makes these demos worth watching
The notable shift is from a single impressive motion to a multi-step job with context: the robot has to understand where it is, choose a path, use its body and place an object in the requested destination. Google also describes a system that can identify task boundaries, track progress and coordinate multiple robots when one machine cannot complete the workflow alone.
The original demonstrations and access details are in Google DeepMind’s release: https://deepmind.google/blog/gemini-robotics-2-brings-whole-body-intelligence-to-robots/. For more physical AI examples, see our Holo4 computer-use demos and the Isaac ROS deployments roundup.
Source published July 30, 2026. Coverage is based on the maker’s announcement and demonstration.
