Google DeepMind has released Gemini Robotics 2, a family of models that can control an entire humanoid robot — walking, crouching, stretching, and manipulating objects — from a single natural language instruction. The system also coordinates multi-robot teams and adapts to entirely new robot bodies in a few hours with fewer than 200 training examples.
The release, announced July 30, represents the most significant leap in DeepMind’s robotics programme to date. Where the previous Gemini Robotics ER 1.5 model controlled upper-body movements for tabletop tasks, version 2 drives the full body — feet to fingertips — and introduces multi-robot collaboration for the first time.
🔍 THE BOTTOM LINE
Gemini Robotics 2 is the first AI model to convincingly control whole humanoid bodies with language commands, not just arms on a bench. The dexterity numbers are still modest — a ziplock bag seals 40% of the time — but the architecture is the breakthrough: one model, any body, hours not months to adapt. DeepMind is calling it a milestone toward “physical AGI,” and for once the framing isn’t unreasonable.
What Gemini Robotics 2 Actually Does
The release includes three distinct models working as a stack:
Gemini Robotics 2 (the VLA) is the vision-language-action model that turns what a robot sees and hears into motor commands. It handles the physical doing — walking, grasping, balancing. This is the model that controls Apptronik’s Apollo 2 humanoid, which can be told “put the watering can into the green bin on the bottom shelf” and execute the full sequence: walk to the table, pick up the can, walk to the shelves, bend down, place it.
Gemini Robotics ER 2 (embodied reasoning) is the high-level brain. It plans multi-step tasks, communicates with humans, tracks progress, and can call external tools like Google Search. It serves as the orchestrator that tells the VLA what to do. ER 2 is available now on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform.
Gemini Robotics On-Device 2 runs locally with no internet connection and can adapt to a completely new robot body — different shape, different sensors, different degrees of freedom — with less than 200 examples and a few hours of training. This addresses what DeepMind calls “one of robotics’ hardest problems”: transferring learned skills between machines.
The dexterity gains are notable. The model can drive Apollo’s five-fingered, 22-joint hand to tie knots and seal ziplock bags, and it can also operate standard two-fingered parallel grippers on Franka Duo platforms for tasks like tight packing.
Impressive, and Still Slow
DeepMind was unusually frank about the limits. Bloomberg reported the system could unscrew a light bulb 92% of the time, but fiddlier tasks lag badly. Trash-bag tying succeeds about 44% of the time. Ziplock sealing: 40%.
The robots are also slow. They pause to think through moves a human makes without a second thought. Kanishka Rao, a DeepMind robotics director, told TNW that true dexterity remains a distant goal and that robots still learn far less efficiently than humans, who adjust after one or two mistakes.
This is a pattern across the field. Rival efforts from robotics foundation-model startups and dexterity work on competing humanoids keep hitting the same wall: the demos dazzle, but the machines are still far from ready for the home.
Multi-Robot Collaboration and Safety
A genuinely new capability is multi-robot teamwork. Gemini Robotics ER 2 can coordinate different types of robots — a wheeled machine and a humanoid, for instance — to split a task neither could complete alone. The reasoning layer assigns sub-tasks, monitors progress, and reallocates if something fails.
On safety, DeepMind introduced ASIMOV-Agentic, a benchmark that tests whether the reasoning model will refuse unsafe commands from the action model and whether it proactively requests human intervention when uncertain. Gemini Robotics ER 2 is described as DeepMind’s safest robotics model to date in safety-constraint following and human-proximity benchmarks — it can detect when a person approaches and bring the robot to a safe stop.
The Hardware Problem Google Cannot Solve
Google makes the software, not the robots. The hardware supply is getting political. As Axios noted, the US just moved to ban future sales of Chinese-made robots on security grounds. Many of the bodies this software might run on are built in China.
Google is working with Western partners — Apptronik, Boston Dynamics, and Agile Robots — plus more than 100 trusted testers. Its rivals are circling the same prize: OpenAI and Nvidia are both building robot models, and everyone is chasing the same idea of one model running any body. This connects to broader moves in physical AI investment and sovereign robotics strategies.
Why This Matters for New Zealand
New Zealand’s robotics sector is small but focused on agricultural and service applications. Whole-body control models that adapt in hours rather than months could lower the barrier for local integrators deploying robots in non-standard environments — orchards, warehouses, aged-care facilities. The on-device model’s ability to run without internet is relevant for rural deployments where connectivity is patchy. The technology is years from commercial deployment, but the trajectory from tabletop demos to whole-body control took DeepMind less than 18 months.
❓ FAQ
Can I buy a robot running Gemini Robotics 2? No. The VLA and On-Device models are available to early-access partners only. The ER 2 reasoning model is available on Google AI Studio for developers. Consumer deployment is years away.
How is this different from what Tesla or Figure showed? Tesla’s Optimus and Figure’s robots are primarily teleoperated or pre-programmed for demos. Gemini Robotics 2 is a learned model that generalises across tasks and robot bodies. DeepMind controls the software layer; its hardware partners build the bodies.
What does “whole-body intelligence” mean? Previous models controlled arms and hands on a tabletop. Gemini Robotics 2 coordinates legs, torso, arms, and hands simultaneously — walking, balancing, and manipulating objects in spaces designed for human movement.
Is the robot safe to be around? DeepMind says ER 2 is its safest model yet, with human-proximity detection and automatic halting. The ASIMOV-Agentic benchmark tests whether the reasoning layer refuses unsafe commands. Real-world safety certification is a separate, longer process.
🔍 THE BOTTOM LINE
Gemini Robotics 2 closes the gap between “robot arms doing tricks” and “robots doing useful work in human spaces” — not all the way, but measurably. The whole-body control, multi-robot collaboration, and hours-fast adaptation are each significant on their own. Together, they signal that the physical AI layer is maturing faster than the hardware it runs on. The next bottleneck won’t be intelligence. It will be bodies, cost, and the political fight over who builds them.