Artificial Intelligence (AI)

Google DeepMind Launches Gemini Robotics 2.0 Suite with Three Models and New Safety Benchmark

Google DeepMind has launched Gemini Robotics 2.0, a three-model suite covering embodied reasoning, robot action generation, and on-device inference, with a new ASIMOV-Agentic safety benchmark and the on-device model adapting to new robot hardware with around 200 examples.

By Daniel Krauss | Edited by Kseniia Klichova Published:
A humanoid robot executing a multi-step manipulation task guided by a three-model AI stack covering embodied reasoning, action generation, and on-device inference, with real-time failure detection and safety halting when humans approach. Photo: Google DeepMind

Google DeepMind has launched Gemini Robotics 2.0, a suite of three AI models designed to give robots the ability to understand instructions, plan multi-step tasks, execute physical actions, and operate safely alongside humans. The release represents the most comprehensive update to Google’s robotics AI platform since it began publishing Gemini-powered robot models, and introduces a new safety evaluation framework alongside improved performance across all three model tiers.

Gemini Robotics ER 2, the embodied reasoning model, is publicly available to developers via the Gemini API and Google AI Studio. Gemini Robotics 2, the action model, and Gemini Robotics On-Device 2, a low-latency offline version, are currently limited to a small group of testers.

The Three-Model Architecture

The suite is structured as a layered stack. Gemini Robotics ER 2 is a vision-language model that processes live video feeds, understands natural language instructions, tracks task progress, and hands off to the action model once it has mapped out what needs to happen. It integrates with the Gemini Live API’s bidirectional streaming endpoint for low-latency orchestration.

Gemini Robotics 2 is the action model – the layer that converts ER 2’s plans into robot movements, generating actions in the same way other generative AI systems produce text or images. The on-device version, Gemini Robotics On-Device 2, runs locally without cloud connectivity for latency-sensitive deployments. Google says the on-device model can adapt to new robot hardware designs with approximately 200 examples or a few hours of movement data.

Performance Improvements

Gemini Robotics ER 2 classifies video frame completeness – tracking how far through a task a robot has progressed – with almost 60% accuracy. That figure is significantly better than the 1.6 release and outperforms comparable visual understanding from competing AI models, though it remains far from reliable for high-stakes deployment without additional safeguards.

Moment-finding – identifying the precise video frame where a critical event occurs, such as when a cup is full or a bolt is tightened – improves substantially, reaching almost 90% accuracy. This enables robots to switch between task steps precisely rather than approximating transitions.

When something goes wrong mid-task, ER 2 can identify the failure in real time and retry the specific failed step without restarting the entire workflow. Video demonstrations show a robot readjusting its hand position when a ball rolls away or a container is moved, continuing the task rather than stopping for human intervention.

Multi-Robot Collaboration

Gemini Robotics 2.0 enables multi-robot collaboration through shared semantic understanding, allowing different robot types to coordinate task handoffs. DeepMind demonstrated this with Apptronik’s Apollo 2 humanoid and a Franka F3 Duo manipulator working together without interference – each robot handling the portion of a workflow suited to its hardware while ER 2 orchestrates the overall sequence.

ASIMOV-Agentic Safety Benchmark

The release introduces ASIMOV-Agentic, a new safety benchmark for physical AI systems that evaluates whether an embodied reasoning model will refuse unsafe tool calls from a VLA model, assess whether a given task can be completed safely, and determine when to call for human assistance rather than proceeding autonomously. The benchmark is publicly available on Hugging Face. DeepMind describes Gemini Robotics ER 2 as its safest robotics model yet, demonstrating the ability to halt a humanoid robot when a human is detected in close proximity and autonomously resume only once the area is clear.

The goal DeepMind scientists describe for the Gemini Robotics program is physical AGI – a generalist robot that can execute any task a human could perform given natural language instructions. The 2.0 release is a step in that direction, with the action model and on-device model representing the components that will determine whether the embodied reasoning advances translate into physical capability at deployment scale.

Disclaimer: RobotsBeat is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, Robotics, technology, software, and digital innovation sectors. These relationships do not influence RobotsBeat's editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

Artificial Intelligence (AI), News, Robots & Robotics

More from RobotsBeat

Exit mobile version