Google DeepMind has launched Gemini Robotics ER 2, its most capable embodied reasoning model for robotics, introducing multi-robot collaboration, continuous video-based task progress classification, and improved spatial reasoning over its predecessor, Gemini Robotics ER 1.6. The model is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform.
Gemini Robotics ER 2 is designed as a high-level brain for robotic systems – handling human conversation, physical world understanding, and multi-step task planning, then handing motor execution to lower-level vision-language-action models. It integrates with the Gemini Live API’s bidirectional streaming endpoint, enabling fluid orchestration without the stop-and-think pauses that have limited prior approaches to high-level robot reasoning.
Multi-Robot Collaboration
The most significant new capability is multi-robot collaboration, allowing robots of different types to communicate via shared semantic understanding and coordinate task handoffs that no single robot could complete alone. DeepMind demonstrated this with Apptronik’s Apollo 2 humanoid and a Franka F3 Duo manipulator arm working together in a shared workspace, with Gemini Robotics ER 2 orchestrating the handoff between them.
The architecture reflects a practical reality of industrial and service robotics deployment: different robot form factors excel in different conditions. A wheeled rover handles flat indoor environments efficiently; a humanoid handles uneven terrain and manipulation tasks better. A model that can coordinate across heterogeneous fleets rather than requiring a single robot to handle every task is closer to how real-world multi-robot systems would operate.
Temporal Intelligence and Task Completion
Knowing when a task is finished – precisely, not approximately – is one of robotics’ hardest unsolved problems. Gemini Robotics ER 2 addresses this through two new capabilities. Continuous progress classification assigns each frame in a live video feed to one of five progress bands, giving robots real-time situational awareness to adjust actions or retry failed steps without restarting complete workflows. Moment-finding identifies the exact video frame where a critical event occurs – when a cup is full, when a bolt is tightened, when a bag is tied – enabling precise task transitions and success verification.
These capabilities allow the model to track its own progress through continuous video feeds rather than relying on static snapshots or human confirmation before proceeding to the next step.
Agentic Task Orchestration
Gemini Robotics ER 2 is built as a physical agent rather than a reactive model. Developers can declare low-level control interfaces – VLA models, navigation APIs, or custom tools – as callable functions, and stream multimodal video, audio, or text directly into the model. It can call Google Search or user-defined functions as needed during task execution. DeepMind demonstrated this with Boston Dynamics’ Spot, using Gemini Robotics ER 2 to orchestrate Spot’s navigation and manipulator APIs for interactive object retrieval.
Safety Improvements
DeepMind describes Gemini Robotics ER 2 as its safest robotics model to date, with significant gains on Safety Instruction Following and Human Proximity benchmarks. The model successfully halts a humanoid robot when a person is detected nearby and autonomously resumes work only once the area is clear. A new safety benchmark evaluates a foundation model’s ability to enforce safety constraints, monitor the environment, assess physical feasibility, and seek human clarification during VLA orchestration. DeepMind published a safety technical report alongside the launch.
Disclaimer: RobotsBeat is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, Robotics, technology, software, and digital innovation sectors. These relationships do not influence RobotsBeat's editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.
