Google DeepMind launches Gemini Robotics 2 with whole-body control and multi-robot coordination
Gemini Robotics 2 enables humanoids to perform dexterous tasks and coordinate with other robots, with on-device inference for rapid adaptation to new embodiments.
Last verified:
The Release
According to DeepMind, Gemini Robotics 2 represents a significant advancement in robotic autonomy, moving beyond pre-programmed or teleoperated systems toward machines capable of reasoning about complex, multi-step tasks in unstructured environments. The release consists of three complementary models designed to address different constraints in robotic deployment: inference speed, adaptability, and coordination complexity.
Gemini Robotics 2: Full-Body Motor Control
The flagship vision-language-action (VLA) model converts visual and language input directly into motor commands, enabling humanoids to perform tasks requiring coordinated whole-body movement. According to DeepMind, Gemini Robotics 2 can orchestrate full-body behaviors—walking, crouching, stretching—while simultaneously controlling manipulators for object interaction. This unified architecture consolidates foot placement, trunk balance, and arm dexterity into a single learned policy, rather than relying on separate sub-systems for locomotion and manipulation as prior generation systems typically do.
Embodied Reasoning and Multi-Robot Teamwork
DeepMind introduced Gemini Robotics ER 2, a vision-language model optimized for planning and reasoning over sequences lasting several minutes. The embodied reasoning variant enables robots to parse natural language instructions, decompose them into multi-step action plans, and communicate progress to human operators. Critically, this model also coordinates behavior across multiple robots, allowing heterogeneous teams to divide labor on complex tasks—a capability that prior robotics models typically did not support.
On-Device Efficiency and Rapid Embodiment Transfer
The on-device variant, Gemini Robotics On-Device 2, runs inference locally on robotic hardware without cloud compute. According to DeepMind, this model achieves fast adaptation to entirely new robot morphologies using only hours of collected data, addressing a longstanding bottleneck in hardware generalization. This capability reduces the time-to-deployment for novel robotic bodies from weeks of retraining to single-digit hours of fine-tuning.
Availability and Developer Access
DeepMind reports that Gemini Robotics ER 2 is available via Google AI Studio and private preview on the Gemini Enterprise Agent Platform. The VLA and on-device models remain in early-access with partner robotics companies, suggesting a staged rollout rather than immediate public release.
Why This Matters
The release signals Google DeepMind’s conviction that foundation model architecture—specifically vision-language models—can unify discrete robotic tasks that previously required hand-engineered controllers. Teams deploying humanoids in logistics, hospitality, or manufacturing will need to evaluate whether Gemini Robotics 2’s generalization across embodiments justifies dependency on a single unified architecture versus modular classical controllers.
The multi-robot coordination capability also reshapes the economics of fleet deployment: if task decomposition and inter-robot communication can be learned rather than scripted, operations teams can scale heterogeneous fleets without proportional increases in programming overhead. The on-device inference constraint is particularly relevant for field deployment, where latency, privacy, and connectivity are non-negotiable requirements. Organizations building robotic systems should monitor early-access partner results to assess whether the hours-to-adapt promise holds under production conditions.
Frequently Asked Questions
What is Gemini Robotics 2?
Gemini Robotics 2 is a vision-language-action (VLA) model that converts visual and language input into motor control commands, enabling robots to perform complex whole-body tasks including walking, manipulation, and object interaction.
Can it work on different robot bodies?
Yes. The on-device variant, Gemini Robotics On-Device 2, adapts to new robot embodiments in just a few hours of data, eliminating the need for extensive retraining.
Does it support multi-robot coordination?
Yes. Gemini Robotics ER 2, the embodied reasoning model, enables robots to communicate with humans, understand their environment, and collaborate with other robots on multi-step tasks.
Where can developers access these models?
Gemini Robotics ER 2 is available on Google AI Studio and in private preview on Gemini Enterprise Agent Platform. The VLA and on-device models are available to early-access partners.