2 September 2026
Google's RT-2 model enabled robots to follow language commands
First reported
Understanding AI ran this on .
- Google released RT-2 in July 2023, a system that let robots understand images and follow verbal instructions to perform physical tasks.
- RT-2 trained on internet knowledge, allowing robots to handle unfamiliar objects and environments without being specifically programmed for each scenario.
- The release prompted investment and startup formation around vision-language-action models, a category combining image recognition, language understanding, and physical action.
How it was covered
Understanding AITimothy B. Lee
Google's July 2023 RT-2 model trained a large multimodal LLM to directly generate robot actions, enabling robots to generalize across objects and scenes using internet knowledge. The breakthrough sparked industry investment and startups, establishing VLA (vision-language-action) models as the standard approach for robot control.