Google LLC has unveiled Robotics Transformer 2 (RT-2), a groundbreaking artificial intelligence model that enables seamless communication between humans and robots. This cutting-edge model transforms spoken words into actionable commands, revolutionizing the way we interact with robotic systems.
RT-2 represents a significant leap forward in the field of robotics as it can learn from both verbal instructions and visual cues, comprehending ideas and concepts to translate them into robotic actions. From picking up objects to triggering various tasks, RT-2 leverages its capacity to grasp information from diverse sources, including words and visuals, to make robots more helpful and efficient.
According to Vincent Vanhoucke, distinguished scientist and head of robotics at Google DeepMind, developing genuinely useful robots has always been an immense challenge. Such robots must possess the ability to handle complex, abstract tasks in diverse environments, even those they have never encountered before.
The new AI system, RT-2, belongs to a novel category called a vision-language-action model (VLM), capable of integrating the semantic meaning of text and visual data. This allows the model to execute intricate tasks by reasoning through a series of instructions. For instance, it can pick up an object and place it in a designated location, like throwing away trash into a bin, or select a snack for a tired individual, such as an energy drink.

Unlike previous AI models like OpenAI LP’s ChatGPT or Google’s Bard, VLMs possess the ability to synthesize both textual and visual data, resulting in a more coherent understanding of complex concepts to accomplish tasks. This presents a fresh set of challenges for robotics engineers, necessitating the formulation of objectives that enable robots to generalize needs based on requests.
RT-2’s capabilities are particularly impressive when dealing with ambiguous scenarios. For example, it can discern what constitutes trash by drawing from a vast corpus of training data, identifying crumpled paper, discarded wrappers, or torn-off straw tips without explicit training for each item.
In contrast to previous approaches that involved complex stacks of systems, RT-2 eliminates this complexity by relying on a single model to handle high-level reasoning and low-level manipulation for robot actions. This seamless integration of functionalities allows robots to respond efficiently to commands and prompts.
Google’s earlier visual model, PaLM-E, played a foundational role in RT-2’s development, enabling robots to make sense of their surroundings and execute sequential tasks through voice commands. Building upon the achievements of the previous model, RT-2 aims to achieve web-scale capabilities, enabling it to handle novel tasks and scenarios it has never encountered before.
In testing, RT-2 demonstrated remarkable performance. It retained the same level of proficiency as its predecessor, RT-1, for tasks within its training data. Additionally, it successfully performed novel tasks or responded to new questions 62% of the time, compared to RT-1’s success rate of 32%.
Vincent Vanhoucke expresses great excitement about RT-2’s potential, not only for advancing AI in robotics but also for creating more versatile and general-purpose robots. While there is still much work to be done to enable helpful robots in human-centered environments, RT-2 provides a promising glimpse into an exciting future for robotics.

