Gemini Robotics 2 aims to broaden robot control

Google Deepmind has introduced Gemini Robotics 2, a vision-language-action model for robots that operate in physical environments. The company says it can support systems from tabletop arms to full-body humanoid robots, while Gemini Robotics ER 2 handles higher-level embodied reasoning.

Gemini Robotics 2 aims to broaden robot control

Google Deepmind has introduced Gemini Robotics 2, a new vision-language-action model designed to help robots connect what they see, what they understand, and what they do. The company describes it as its most advanced VLA model yet and positions it as an intelligence layer for adaptive robots.

What Gemini Robotics 2 is designed to do

Gemini Robotics 2 is built around a vision-language-action approach. In plain terms, that means the model combines image recognition, language processing, and action control so a robot can operate in a physical environment.

That combination matters because robots need more than one kind of intelligence to be useful outside narrow, fixed routines. They must interpret visual information, understand instructions or context, and translate that understanding into movement.

According to Deepmind, Gemini Robotics 2 can control a range of systems, from tabletop arms to full-body humanoid robots. That range is central to the announcement: the model is not being described as a tool for only one robot shape or one narrow task category.

From fine movement to full-body control

Deepmind says Gemini Robotics 2 can manage full-body movement and perform fine motor tasks. Those are different kinds of robotic control problems, and placing them under the same model family suggests a push toward more general robot behavior.

Fine motor tasks require careful action at a small scale. Full-body movement involves coordinating a larger physical system. A model that can address both is meant to support robots that do more than repeat a single preprogrammed motion.

The company also says Gemini Robotics 2 can coordinate multiple robots. That points to a broader role than controlling one machine in isolation. In physical environments, useful robotic systems may need to work together, divide actions, or respond to shared surroundings.

Why VLA models matter for physical AI

VLA models are important because they connect perception and action. A robot that recognizes objects but cannot act on that recognition is limited. A robot that can move but cannot understand what it sees or what it is being asked to do is also limited.

Gemini Robotics 2 sits at that intersection. It is presented as a model that helps robots operate in the real world by linking images, language, and movement into one control system.

The source description uses the phrase "intelligence layer" for this role. That is a useful way to understand the model: it is not simply a visual detector or a language interface, but a system intended to guide robotic behavior.

  • Vision helps the robot process physical surroundings.
  • Language helps the system handle instructions and meaning.
  • Action control turns perception and understanding into movement.

Within the limits of the announcement, the key point is not that every robot will immediately gain these capabilities. The point is that Deepmind is offering a model architecture intended to support robots of many forms, including arms and humanoids.

Gemini Robotics ER 2 adds embodied reasoning

Alongside Gemini Robotics 2, Google Deepmind introduced Gemini Robotics ER 2. This model is designed for "embodied reasoning," a term the source defines as understanding the physical world and deciding which actions to take based on that information.

That makes ER 2 a higher-level control system. While Gemini Robotics 2 focuses on vision-language-action capabilities, ER 2 is described as helping robots reason about the physical context before acting.

Deepmind says ER 2 replaces Gemini Robotics ER 1.6, released in April. The new model is available in Google AI Studio.

For developers, access paths differ across the two announcements. Gemini Robotics ER 2 is available in Google AI Studio, while developers can apply for early access to Gemini Robotics 2 through the waitlist.

The practical takeaway

The announcement shows Google Deepmind continuing to organize robotics around models that combine perception, language, reasoning, and action. Gemini Robotics 2 is framed as the main VLA model for controlling robots across different body types, while Gemini Robotics ER 2 is framed around embodied reasoning.

For now, the practical details in the source are limited to capabilities, availability, and positioning. The confirmed facts are clear: Gemini Robotics 2 is aimed at adaptive robots, can control systems from tabletop arms to full-body humanoid robots, and is available to developers through an early access waitlist. Gemini Robotics ER 2 replaces Gemini Robotics ER 1.6 and is available in Google AI Studio.