Google DeepMind has released a new version of its artificial intelligence model Gemini that can control a range of robots, including humanoids capable of dextrous tasks such as screwing in lightbulbs and tying trash bags, according to WIRED. The system, called Gemini Robotics 2, combines several different AI models into a single system to enable robots to make sense of their surroundings and act within them.
How Gemini Robotics 2 Works
Gemini Robotics 2 integrates a vision language model (VLM), which understands images and video and can communicate with humans and reason how to perform different tasks, along with two vision language action (VLA) models. According to WIRED, the VLA models are trained to understand how to move in physical space and control the robot's full-body movement as well as the movements of grippers or hands.
Google DeepMind trained the model using a mix of human teleoperation, video examples, and simulations. WIRED notes that it is not yet possible for AI models to perform a wide range of complex tasks without specific training.
Demonstration: Tidying Shelves with Apptronik's Apollo 2
In video demonstrations shared ahead of the release, the company showed several different robots performing complex tasks autonomously using the amalgamated model, according to WIRED. One demo featured Apptronik's Apollo 2 robot using hands from a company called Sharpa to tidy shelves.
Safety and Risks
Giving frontier AI models access to robots that can wander workplaces or homes and manipulate objects comes with risks. WIRED reports that previous research has shown using frontier AI to control robots can produce unexpected and sometimes dangerous behavior. The safety question is even more pressing when robots are placed in many different situations with uncertainty.
Carolina Parada, head of robotics at Google DeepMind, told WIRED: "The safety question is even more pressing because you're putting them in a lot of other situations. There's a lot of uncertainty that will show up, and so you want to be able to understand the safety question more deeply."
Parada says Google takes a multi-layered approach to safety, with guardrails applied on each model layer. The company is also introducing ASIMOV-Agentic, a new benchmark for measuring the safety of various AI systems collaborating to control a robot. The benchmark detects whether a command will result in a harmful or uncertain outcome.
Google's Robotics Ambitions
Although Anthropic and OpenAI have taken a lead with chatbots and AI coding tools, Google has a stronger track record in robotics research, according to WIRED. The search giant previously partnered with Boston Dynamics, a leader in legged robots, to provide the brains for those machines. The release is another sign that Google is betting AI will need to break free from the digital realm to realize its full potential.
Parada stated: "It's another milestone in our path towards really getting towards what we call like physical AGI, which means we get a robot to do anything that a human can."
Google's CEO, Demis Hassabis, previously told WIRED that he hopes to develop an AI operating system for many different robots, similar to the Android operating system for smartphones.
Implications for Enterprise Automation
For enterprise technology decision-makers, particularly those in logistics and supply chain, the ability to control humanoid robots with a unified AI model opens opportunities for automating tasks such as shelf-tidying, light assembly, and packaging. However, the safety concerns highlighted by Google DeepMind underscore the need for rigorous testing and guardrails before deployment in commercial environments. The introduction of the ASIMOV-Agentic benchmark provides a framework for evaluating the safety of multi-model robot control systems, which could become a standard for enterprise adoption.