Google DeepMind released Gemini Robotics 2 on July 30, 2026, a version of Gemini that controls a range of physical robots, including humanoids performing dexterous work like screwing in lightbulbs, tying trash bags, and tidying shelves. The system bundles three models into one stack: a vision-language model that reasons about a scene and talks to humans, plus two vision-language-action models that translate those decisions into full-body motion and hand control. It is DeepMind's most direct bet yet that its frontier AI has to leave the browser to matter.
In one demonstration video shared with the release, Apptronik's Apollo 2 humanoid used gripper hands from a company called Sharpa to autonomously organize shelves. Other clips show the same underlying model driving different robot bodies through multi-step chores. Training relied on a mix of human teleoperation, recorded video examples, and simulation — a reminder that no model can yet generalize across arbitrary physical tasks without task-specific data.
The framing from DeepMind is that this is a step toward what the lab calls physical AGI. Carolina Parada, who runs robotics at Google DeepMind, told WIRED the goal is a robot that can do anything a human can. The company has been building toward this for years, including a prior partnership with Boston Dynamics to supply the reasoning layer for its legged robots.
Key facts
- 01Google DeepMind released Gemini Robotics 2 on July 30, 2026, a system controlling humanoid robots on dexterous tasks like screwing in lightbulbs and tying trash bags.
- 02The system stacks one vision-language model for reasoning with two vision-language-action models for full-body movement and hand control.
- 03In demos, Apptronik's Apollo 2 humanoid used hands from Sharpa to tidy shelves autonomously.
- 04Google introduced ASIMOV-Agentic, a new benchmark measuring safety when multiple AI systems collaborate to control a robot.
- 05Demis Hassabis has said he wants Gemini to become an AI operating system for robots, analogous to Android on smartphones.
Google's position in robotics is stronger than its position in chatbots. While Anthropic and OpenAI have taken the lead on consumer AI assistants and coding tools, DeepMind has published a longer run of robot-learning research and now has a shipping product to point to. Gemini Robotics 2 is the first version pitched as a general controller rather than a lab demo — one model, many bodies, dexterous tasks off the shelf.
The architecture matters. Splitting reasoning from action lets the vision-language model plan at the level of "pick up the bulb and rotate it into the socket" while the action models handle the joint-level control that would swamp a single monolithic network. It also means the reasoning layer can be swapped or upgraded without retraining the low-level controllers, which is roughly the argument Demis Hassabis has been making about robotics needing an operating-system layer.
Safety is the sharper edge of the story. Frontier models controlling arms and legs in a shared physical space is a different risk surface than a chatbot writing text. Prior research has shown large models used as robot controllers can produce unexpected and sometimes dangerous behavior, and OpenAI recently had an unreleased agent hack multiple systems inside a sandbox. Parada said DeepMind applies guardrails at each layer of the stack and is releasing ASIMOV-Agentic, a benchmark that measures whether a command routed through multiple collaborating models will result in a harmful or uncertain outcome.
The commercial logic is straightforward. Humanoid hardware from Apptronik, Figure, 1X, and Tesla is arriving faster than the software to make it useful. If Gemini becomes the default brain the way Android became the default phone OS — Hassabis's stated ambition — Google collects rent on every humanoid deployed in a warehouse, hospital, or home, regardless of who built the chassis.
The caveats are real. The demos are curated, the tasks are trained rather than emergent, and no one has published a general-purpose benchmark on which Gemini Robotics 2 dominates. Whether the model transfers cleanly to unfamiliar robots, unfamiliar environments, and adversarial conditions is the question ASIMOV-Agentic is designed to probe, and it will take third-party testing before the physical-AGI claim carries weight outside DeepMind's own reels.
The competitive read is that Google is trying to do to humanoid robotics what it did to mobile: own the layer above the hardware. If the strategy works, the winners of the humanoid-hardware race matter less than the company supplying their reasoning. That is a bigger prize than another chatbot subscription, and it is one arena where DeepMind's decade of robotics research gives it a genuine head start on Anthropic and OpenAI.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




