Google DeepMind’s new AI model can control a robot’s entire body

WorkAI.TV Editorial Desk
3 Min Read

Share with your CTO

Google DeepMind is betting that whole-body robot control, not just arm manipulation, is the unlock the industry has been waiting for, and it’s shipping Gemini Robotics 2 to prove it. The new model extends coordination from fingertips to feet, enabling humanoid robots like Apptronik’s Apollo 2 to crouch, walk, and retrieve objects in a single continuous motion. The companion Gemini Robotics ER 2 vision-language model now handles multi-step tasks, understands task boundaries, and coordinates across heterogeneous robot fleets without cloud dependency on the on-device variant.

What this means for your business

The gap between “robot that moves an arm” and “robot that navigates a real environment” has been the structural ceiling on warehouse, logistics, and manufacturing automation pilots. If your organization has run humanoid robot evaluations and hit that ceiling, Gemini Robotics 2 is directly aimed at the failure mode you documented. If you haven’t started, the competitive timeline just compressed: Google is positioning whole-body coordination as a shipping capability, not a research milestone.

The multi-robot orchestration feature deserves more attention than the headline whole-body story. A single robot completing a task is a demo. One robot issuing instructions to a different-form-factor robot to complete a parallel subtask is the beginning of a composable physical automation layer, the idea that robots of different designs can be assigned roles in a workflow the way software microservices are. That architectural pattern, if it holds outside controlled demos, changes the procurement calculus from “buy a robot” to “design a robot workflow,” which is a fundamentally different vendor conversation.

The falsification condition here is movement speed. DeepMind explicitly flags it as an open limitation, and speed isn’t cosmetic in industrial settings where cycle time drives ROI. If Gemini Robotics 2 deployments in 2025 can’t match even 60 percent of human task throughput in real facilities, the whole-body coordination story stays a research showcase. Watch Apptronik’s Apollo 2 commercial deployment numbers, not the YouTube demos, as the leading indicator of whether this model has crossed the operational threshold.

Concept deep-dive: Vision-language model for robotics

A vision-language model (VLM) in robotics combines a camera feed with natural-language instructions to decide what a robot should do next, roughly like giving a robot both eyes and the ability to read a task description simultaneously. Gemini Robotics ER 2 is Google’s VLM layer sitting above the motion controller. The business relevance is that VLMs reduce the need for hard-coded task programming: a robot can receive a new instruction in plain language and reason about how to execute it in an unfamiliar environment.

Based on reporting from Google DeepMind’s new AI model can control a robot’s entire body, originally published 2026-07-30 13:18:00.

TAGGED:
Share This Article