Products

Google DeepMind’s Gemini Robotics 2 brings whole-body control to humanoids

The new foundation model suite coordinates entire robots from feet to fingertips, adapts to new hardware in hours, and introduces multi-robot teamwork.

Editorial·26 Aug 2026
Google DeepMind’s Gemini Robotics 2 brings whole-body control to humanoids

Google DeepMind has unveiled the second generation of its robotics foundation models, a suite headlined by a capability the company describes as “intelligent whole-body control.” Announced on July 30, 2026, Gemini Robotics 2 moves beyond the tabletop manipulation focus of its predecessor to coordinate an entire humanoid robot “from feet to fingertips.” The flagship vision-language-action (VLA) model is accompanied by two companion systems: Gemini Robotics ER 2, an embodied-reasoning model that acts as a high-level planner, and Gemini Robotics On-Device 2, a compact VLA designed to run locally on robot hardware. The announcement was authored by Carolina Parada of Google DeepMind.

The announcement signals a decisive shift in the trajectory of physical AI. For executives and engineers tracking the robotics market, the most consequential change is not a single benchmark improvement but a structural one: a single model checkpoint can now control multiple robot bodies with very different shapes, sensors, and degrees of freedom, and can be adapted to new embodiments in hours rather than months. That has direct implications for the cost and speed of deploying robots in logistics, manufacturing, and service environments, where hardware diversity has long forced expensive, per-robot software development. DeepMind’s move also arrives as venture capital continues to pour into humanoid robotics startups, while established industrial automation vendors race to integrate foundation models into their product lines.

Whole-body control and dexterous manipulation

The headline demonstration pairs Gemini Robotics 2 with Apptronik’s Apollo 2 humanoid, equipped with a five-fingered, 22-degree-of-freedom SharpaWave hand. In videos released alongside the announcement, the robot ties knots and seals ziplock bags — tasks that require coordinated use of the torso, arms, wrists, and individual fingers. DeepMind also showed the model controlling a Franka Duo platform with a two-fingered parallel gripper to perform tight packing tasks, illustrating that the same underlying model can switch between drastically different manipulator types without retraining from scratch.

Carolina Parada framed the advance as a move from “narrow, task-specific” robotics toward models that can reason about their entire body. The company reports medium-to-high success rates on whole-body and gripper-based tasks. However, DeepMind is explicit about the limits: multi-finger dexterous manipulation “remains challenging,” and movement speed still needs advancement. The admission is notable for its candor in a field often dominated by polished demo videos, and it provides a useful calibration for procurement teams evaluating near-term deployment feasibility. The gap between tying a knot in a controlled lab setting and doing so reliably on a moving production line remains wide, and DeepMind’s own language signals that full dexterity at human speed is not yet solved.

Fast adaptation and multi-robot teamwork

One of the most commercially relevant claims in the announcement concerns adaptation speed. DeepMind states that Gemini Robotics 2 can be adapted to new robot embodiments in “a few hours,” typically with fewer than 200 examples. The company demonstrated this on three additional platforms — Dexmate, SO101, and Trossen — which differ substantially in form factor, sensor suites, and actuation. This capability addresses a persistent bottleneck in industrial robotics: the time and expertise required to port a control policy from one robot model to another. In traditional automation, switching from a six-axis arm to a mobile manipulator or a humanoid often means rebuilding the software stack from the ground up. A few hours of adaptation data, by contrast, could let a warehouse operator test a new robot form factor in a single shift.

The suite also introduces multi-robot collaboration. Different robot types can now communicate and work as a team under the orchestration of Gemini Robotics ER 2, which DeepMind describes as the robot’s “high-level brain.” ER 2 executes multi-step sequences lasting several minutes and involving hundreds of decisions, with self-correction and progress tracking. That is a meaningful extension of the planning horizon compared with earlier VLA models, which typically handled short, single-task episodes. For warehouse and factory operators, the combination of fast embodiment adaptation and heterogeneous fleet coordination suggests a future where a single software stack governs a mixed inventory of robots — from fixed arms to mobile manipulators to humanoids — without per-model engineering teams.

Safety benchmarks and enterprise availability

DeepMind paired the model release with ASIMOV-Agentic, a new benchmark designed to measure an agent’s ability to refuse unsafe tool calls and request human intervention when uncertain. The company calls ER 2 its “safest robotics model to date” on safety-constraint and human-proximity benchmarks. This is not a peripheral addition. As regulators in the European Union, the United States, and elsewhere move toward mandatory safety and liability frameworks for autonomous systems, benchmark-driven safety claims are becoming a prerequisite for enterprise procurement and insurance coverage. A model that can demonstrate consistent refusal of unsafe actions — and do so on a published, reproducible benchmark — gives risk committees a concrete metric to evaluate, rather than relying on vendor assurances.

On availability, the rollout is deliberately staged. Gemini Robotics ER 2 is live on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform. The flagship VLA and the On-Device model are currently open only to early-access partners. DeepMind has not published a timeline for broader release, nor has it disclosed pricing. The staged approach mirrors patterns seen in earlier foundation model launches, where enterprise platform access precedes open or API-based distribution. For now, the practical audience is limited to companies with existing Google Cloud relationships or those willing to join an early-access program with unspecified terms.

Market context and what comes next

The announcement lands in a crowded and rapidly consolidating field. Humanoid robotics startups have raised billions in venture capital over the past two years, while established industrial automation vendors are racing to integrate foundation models into their product lines. DeepMind’s differentiation is its combination of model scale, Google Cloud distribution, and now a safety benchmark that speaks directly to enterprise risk committees. The inclusion of Apptronik — a company with active pilot programs in logistics and automotive manufacturing — signals an intent to move beyond research demos into operational settings. Apptronik’s Apollo 2 platform gives DeepMind a credible hardware partner with real-world deployment experience, rather than a lab-only prototype.

The unresolved questions are substantial. DeepMind has not disclosed the compute requirements for running Gemini Robotics 2 in production, nor the latency characteristics of the full-body control loop. The On-Device 2 model suggests an awareness that cloud-only inference is impractical for many real-time manipulation tasks, but its performance relative to the flagship model remains unspecified. And while the ASIMOV-Agentic benchmark addresses refusal behavior, it does not yet cover the full spectrum of physical safety risks, such as force limits during contact with humans or failure modes under sensor degradation. Those gaps matter for any deployment in environments where robots work alongside people, not just behind cages.

Still, the direction of travel is clear. Whole-body control, fast embodiment adaptation, and multi-robot orchestration are the capabilities that separate general-purpose robotics from the brittle, single-task automation of the past. Google DeepMind has staked a credible claim to all three in a single release. Whether that claim survives contact with real warehouses, factory floors, and hospital corridors will depend less on demo videos than on the unglamorous work of reliability engineering, safety validation, and enterprise integration that now lies ahead. The next 12 to 18 months will show whether early-access partners can translate lab demonstrations into measurable operational results, and whether DeepMind can close the dexterity and speed gaps it has openly acknowledged.

#robotics #Google DeepMind #AI models #automation

Newsletter

Get the AI news that matters

One short brief with the day's most important AI stories — written for professionals.

We send a confirmation link. No spam. Unsubscribe anytime.

WhatsApp