DevOps / SRE / Platform · 31.07.2026, 16:04 UTC
Gemini Robotics 2 brings us one step closer to physical AGI
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 31.07.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
This week, Google DeepMind revealed Gemini Robotics 2, an intelligence layer comprising three new models to power more adaptable physical AI. Together, Google says these models will give robots more dexterous, full-body control to work together and complete a wide range of multi-step tasks.
Google announced the vision-language-action (VLA) model, Gemini Robotics 2, on Thursday, and converts vision and language inputs into motor control so robots can flex from feet to fingertips with enough dexterity in hands and grippers to complete delicate tasks, like closing a Ziploc bag. The lightweight version, Gemini Robotics On-Device 2, runs locally so robotic applications can keep running even without internet connectivity.
Meanwhile, the embodied reasoning (ER) model, Gemini Robotics ER 2, lets robots understand their surroundings and communicate with humans so they can devise plans to carry out multi-step tasks — think emptying a dishwasher and putting items away.
Chris Matthieu, VP of the developer ecosystem at RealSense, tells The New Stack this is what it’ll take for robots to graduate from simple, isolated actions to real-world physical assistance:
“The hard part isn’t making the first decision — it’s recovering from the hundredth when the world has changed. Doors are closed, objects get moved, people walk into the scene, batteries drain, and sensors become partially occluded.”
Full-body control and greater dexterity
According to Google, combining VLA and ER models means humanoids can go further to complete a range of tasks, literally.
Where its previous Gemini …