DeepMind’s newest robotics release gives humanoid machines coordinated control of their whole frame, from ankles to fingertips, rather than the upper-body-only skills of prior versions.
Gemini Robotics 2 is a vision-language-action model whose single trained checkpoint steered three different robot platforms in demonstrations, including Apptronik’s Apollo 2 equipped with two styles of five-fingered hands. Google DeepMind showed the system sealing a Ziploc bag, tying a trash bag and unscrewing a lightbulb, while conceding that fine multi-finger work and movement speed remain rough edges.
The lab frames the release as groundwork for chores that need the whole body working together, like clearing a cluttered room.
DeepMind also refreshed the models that surround the main release. The embodied reasoning system, Gemini Robotics ER 2, gained sharper scene understanding, instruction parsing and multi-robot coordination. And the on-device variant, built to run without cloud access, now adjusts more rapidly to machines with very different shapes, sensors and joint configurations.
The release lands in a heated robotics race, with labs betting that broad, embodiment-agnostic training data will eventually let one model pilot almost any robot.