From Feet to Fingertips: Google DeepMind’s Gemini Robotics 2 Gives Humanoid Robots Full-Body Control

Reading Time: 5 minutes

Google DeepMind's Gemini Robotics 2 expands AI-powered humanoid robot control from the upper body to the entire body, enabling coordinated walking, crouching, and fine manipulation. Demonstrated on Apptronik's Apollo 2, the model marks a significant step toward general-purpose humanoid robots capable of real-world physical tasks.

The Next Frontier in Humanoid Robotics Has Arrived

For years, the challenge of getting a humanoid robot to move like a human has been split into discrete, frustrating sub-problems: balance the torso, coordinate the arms, manage the legs — and hope the whole system doesn’t collapse into an awkward, jerky mess when any of those sub-systems conflict with one another. Google DeepMind appears to have taken a significant step toward solving that fragmentation problem with the release of Gemini Robotics 2.

As reported by The Verge (https://www.theverge.com/tech/973276/google-deepmind-gemini-robotics-2-whole-body), the new model can now “control entire humanoid robots,” supporting what DeepMind describes as “whole-body motions” that span from a robot’s feet all the way up to its fingertips. That is a meaningful leap from its predecessor, which was limited to controlling only a humanoid robot’s upper body.

What Changed Between Gemini Robotics and Gemini Robotics 2

The original Gemini Robotics model was already a notable achievement. It demonstrated that a large AI model could be grounded in physical reality — learning to perceive an environment, reason about objects, and instruct a robot’s arms and hands to carry out dexterous tasks. But restricting control to the upper body left a significant gap. A robot that cannot integrate its leg movements, crouching posture, and weight distribution with its arm actions is fundamentally limited in the kinds of real-world tasks it can handle.

Gemini Robotics 2 closes that gap. According to Google DeepMind’s announcement, the model enables humanoid robots to walk, crouch, stretch, and manipulate objects — all as part of a unified, coordinated motion system. This is sometimes called “whole-body control” in robotics research, and it is considered one of the harder open problems in the field because it requires the AI to reason simultaneously about dozens of degrees of freedom across an entire mechanical body.

Think of it this way: when a human bends over to pick up a heavy watering can, they do not just extend their arms. They shift their centre of gravity, bend at the knees and waist, adjust the tension in their core, and then use their hands and fingers to grip the object securely. Replicating that kind of seamless, whole-body coordination in a robot — and doing so using a single AI model rather than a patchwork of separate controllers — is exactly the problem Gemini Robotics 2 is targeting.

Apptronik’s Apollo 2: The Robot Putting It to the Test

Google DeepMind has shared videos demonstrating Gemini Robotics 2 in action on Apptronik’s Apollo 2 humanoid robot. The demonstrations are telling. In one clip, the Apollo 2 robot bends over to pick up a watering can — a task that requires coordinated lower-body and upper-body movement. In another, the robot locates and retrieves specific items from a shelf, which demands not just physical dexterity but also visual reasoning and object identification.

Apptronik is one of a growing number of humanoid robotics companies that has been partnering with AI labs to bring general-purpose robot intelligence to life. Apollo 2 is a commercially oriented humanoid platform, and its collaboration with Google DeepMind signals that Gemini Robotics 2 is not purely a research curiosity — it is being tested on hardware that is intended for real-world deployment.

The shelf-retrieval task is particularly worth paying attention to. Finding a specific item on a shelf — among other items, under variable lighting, with potential occlusion — requires the model to integrate vision, language understanding (knowing what object to find), and physical action planning. That the robot can accomplish this while also managing its full-body balance and movement suggests that Gemini Robotics 2 is functioning as a genuinely integrated system rather than a collection of bolted-together modules.

Why Whole-Body Control Matters for Real-World AI Deployment

If you have been following the humanoid robot space — companies like Figure, 1X, Boston Dynamics, and Agility Robotics, among others — you will know that most demonstrations until recently have tended to show robots performing tasks in highly controlled environments, often with limited locomotion. The robots stand in place, or they walk on flat surfaces, but rarely do they combine dynamic locomotion with fine manipulation in the same fluid motion sequence.

Whole-body control is the capability that bridges that gap. It is what would allow a robot to, say, walk across a warehouse floor, crouch down to reach a low shelf, pick up a fragile package, stand back up, and carry it to another location — all without stopping to recalibrate between each phase of the task. That kind of continuous, integrated movement is essential for robots to be genuinely useful in logistics, manufacturing, healthcare support, or domestic assistance settings.

For India’s growing manufacturing and warehousing sector — where labour costs are rising and operational efficiency is increasingly tied to automation — the direction that models like Gemini Robotics 2 are pointing could have significant long-term implications. While commercial deployment of humanoid robots at scale remains years away and the cost of such systems would currently run into tens of lakhs to crores of rupees per unit, the foundational AI capabilities being demonstrated here are what will eventually make that deployment viable.

The Broader Context: Gemini as a Physical-World Model

It is worth stepping back to appreciate what Google DeepMind is attempting with the Gemini Robotics line. The Gemini model family was originally conceived as a multimodal AI — one that could understand text, images, audio, and video. Extending it into robotics means grounding that multimodal understanding in the physical world, where decisions have real consequences and errors cannot simply be regenerated like a bad sentence from a chatbot.

This approach — using a single large foundation model to power both perception and action in a robot — is philosophically different from older robotics paradigms where perception, planning, and control were handled by separate, hand-engineered systems. The bet being made by DeepMind, and by competitors like OpenAI (with its investments in robotics) and Physical Intelligence (pi), is that scaling up AI models and training them on diverse robotic data will eventually yield systems that generalise across environments and tasks in the way that large language models generalised across language tasks.

Gemini Robotics 2’s extension to whole-body control is a concrete step in that direction. It suggests that the model is learning increasingly rich representations of physical interaction — not just “how do I grasp this object” but “how do I orient my entire body to accomplish this goal efficiently and safely.”

What Remains Unknown

The source article on The Verge notes that DeepMind’s full announcement contains additional details, though the publicly available summary is somewhat limited. Several important questions remain open. How does the model perform outside of curated demonstration environments? What is the training data pipeline that enables whole-body coordination — is it primarily simulation, real-world teleoperation, or a combination? How quickly can the model adapt to novel tasks or unfamiliar objects without additional fine-tuning?

These are not criticisms so much as the natural frontier questions that accompany any significant capability announcement in AI robotics. The field moves fast, and demonstration videos — however impressive — are only one data point. Independent evaluation across diverse real-world conditions will ultimately determine how transformative this capability proves to be.

The Takeaway

Gemini Robotics 2 represents a qualitatively meaningful upgrade in what AI can do with a humanoid robot’s body. By extending unified model control from the upper body to the full kinematic chain — feet, legs, torso, arms, and hands — Google DeepMind is addressing one of the core bottlenecks that has kept humanoid robots from performing the kinds of fluid, integrated physical tasks that humans take for granted.

The demonstrations on Apptronik’s Apollo 2, from bending to retrieve a watering can to identifying and pulling specific items off shelves, offer a glimpse of what becomes possible when a single AI model can reason about and coordinate an entire body. Whether Gemini Robotics 2 delivers on that promise at scale and in uncontrolled environments is the next question the field will be watching closely.

Related stories