A telling moment from2023

In Google DeepMind’s RT-2 research, a robot was asked which available object could help hammer a nail. The published example shows it selecting a rock. The interesting element is the task description: it asks for a useful object without simply naming the object to pick up.

The image is historical research from July2023. It does not show a robot completing a carpentry job. The visible action is choosing and lifting an object, which is a narrower but revealing connection between a language instruction and physical behavior.

What changed from RT-1

Google’s December2022 RT-1 work trained a transformer on robot demonstrations, mapping camera observations and language instructions to actions. RT-2 built on that robotics data while adapting larger vision-language models trained on web information. Its research explored how broader concepts could transfer into robot control.

The useful comparison is where knowledge comes from. Demonstrations teach physical behavior in recorded situations. Broader language and image learning can help relate an unfamiliar instruction to an available object. Neither by itself proves a robot will handle every real-world situation.

How to watch today’s robot demos

Separate three questions: what did the robot recognize, what action did it choose, and what physical operation did it finish? A clip can be impressive on the first two while leaving the third untested. A longer task can also contain several opportunities for error after a correct first choice.

That viewing habit makes historical research useful today. Instead of treating each new robot clip as a single score, compare the same capability across demonstrations and keep the setting visible. Here, the specific progress is connecting concepts to object selection—not a claim of general household autonomy.

Primary sources

https://deepmind.google/blog/rt-2-new-model-translates-vision-and-language-into-action/

https://research.google/blog/rt-1-robotics-transformer-for-real-world-control-at-scale/

Explore the original source ↗

Source published 2022 / 2023. Coverage is based on the maker’s announcement and demonstration.