Google DeepMind released Gemini Robotics ER 2 on July 30, an embodied reasoning model that watches continuous video, tracks its own task progress, and orchestrates other robots to finish jobs a single machine cannot. On moment-finding benchmarks the model hits 91.3% accuracy with a 0.96-second mean absolute distance, and it does so at 4x the execution speed of larger competing models. Developers can access it now through the Gemini API, Google AI Studio, and a private preview on the Gemini Enterprise Agent Platform.
The pitch is that ER 2 is not the model moving the robot's joints. It is the planner sitting on top, calling vision-language-action models and hardware APIs as tools while the robot keeps acting. That split is what lets the model reason about what comes next without the stop-and-think pauses that have made previous agentic robot demos look stilted.
“Think of Gemini Robotics ER 2 as a high-level brain for robots. It allows robots to chat with humans, understand the physical world, and plan multi-step tasks.”— Steven Hansen, Senior Staff Software Engineer, Google DeepMind
The gains over Gemini Robotics ER 1.6, released earlier this year, land in three places: video understanding, tool orchestration, and multi-robot handoff. ER 2 classifies progress on a live video feed by binning each frame into one of five completion bands — 0-20%, 20-40%, 40-60%, 60-80%, 80-100% — and posts 57.4% accuracy on that task, which Google DeepMind says outperforms both the prior generation and unnamed frontier competitors. The point is situational awareness: a robot that knows it is 60% through tightening a bulb can retry a failed twist without restarting the workflow.
Key facts
- 01Gemini Robotics ER 2 hits 91.3% accuracy and 0.96s mean absolute distance on moment-finding tasks.
- 02The model runs at 4x the execution speed of larger competing models on the same benchmarks.
- 03ER 2 scores 57.4% accuracy on progress classification, grading video frames across five completion bands.
- 04Available today via the Gemini API and Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.
- 05Demoed orchestrating Boston Dynamics' Spot and coordinating Apptronik's Apollo 2 with a Franka F3 Duo arm.
Moment-finding is the sharper number. Given a video, the model has to pick the exact frame where a critical event happens — the instant to stop pouring coffee, the instant a bag is fully tied. The 91.3% accuracy and sub-second mean error put ER 2 within range of much larger models while running fast enough for real-world control loops. Google DeepMind frames the 4x speed edge as the actual gating factor for safety in physical deployment.
The multi-robot demos are where the orchestration story gets concrete. Google DeepMind paired ER 2 with Spot from Boston Dynamics to run a natural-language fetch task, using the model to call Spot's navigation and manipulator APIs. A separate demo has ER 2 coordinating Apptronik's Apollo 2 humanoid with a Franka F3 Duo arm, with both robots sharing semantic context to hand off subtasks. This builds directly on Google DeepMind's Gemini Robotics 2 release covered here previously, which focused on humanoid dexterity rather than the planner layer.
“One of robotics' hardest challenges is knowing when a task is done.”— Peng Xu, Staff Software Engineer, Google DeepMind
Spatial reasoning also gets a lift. ER 2 was tested on 10 different types of instruments — including digital displays, linear scales, rulers, and liquid thermometers — extending beyond the circular dials that dominated prior benchmarks. Success and failure detection now runs on raw video instead of static snapshots, which is what lets the model catch mid-execution problems like spills or slips instead of only noticing them after the fact.
Safety got its own benchmark. Google DeepMind is introducing an evaluation for how well a foundation model behaves as a VLA orchestrator, testing whether it enforces physical constraints, monitors the environment, and asks for human clarification when it should. In one test, ER 2 halted a humanoid robot when a person entered its workspace and resumed autonomously only after the area cleared. The company published a separate safety technical report alongside the release.
The caveats are real. Google DeepMind's benchmark comparisons name ER 1.6 explicitly but describe the other frontier models generically, which makes head-to-head claims hard to verify independently. The 57.4% progress-classification score, while a lead, is still well short of what production deployment in a factory or hospital would require without a human in the loop. And the multi-robot orchestration demos are curated — the Apollo 2 and Spot workflows shown are simpler than the messy handoffs a warehouse floor would generate.
The strategic read is that Google DeepMind is trying to become the operating system for other people's robots rather than the robot maker itself. By exposing ER 2 through the Gemini API and pushing it into the Enterprise Agent Platform, the company is betting that hardware vendors — Boston Dynamics, Apptronik, Franka, and whoever comes next — will plug in rather than build their own planning stacks. If that bet holds, the model becomes the layer every serious commercial robot calls, which is a far bigger business than selling any single humanoid.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




