Gemini Robotics 2: read the success rates, not the demo
Google DeepMind shipped three Gemini Robotics 2 models. Only one is open to developers, and the published success rates are the number that matters.
Google DeepMind released Gemini Robotics 2 today — three models covering whole-body humanoid control, embodied reasoning, and on-device inference. The demo videos are impressive. The success rate tables published alongside them are more useful, and they say something different.
What actually happened
Per DeepMind's announcement, three models shipped together:
Gemini Robotics 2 is the vision-language-action model — it converts camera and language input into motor control, driving full humanoids from feet to fingertips as well as bi-arm setups. Gemini Robotics ER 2 is the embodied reasoning model: it reads a physical environment, plans multi-step tasks running several minutes, and coordinates multiple robots working together. Gemini Robotics On-Device 2 runs locally with no internet connection and adapts to an entirely new robot body with a few hours of data.
Access is tiered, and the tiers matter more than the capabilities. ER 2 is available now in Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. The other two — the VLA and the on-device model — are early-access partners only, behind a sign-up form.
Now the numbers. DeepMind's own reporting puts whole-body manipulation success at 45.7–76.3%, gripper dexterity at 74.2–89.6%, and multi-finger dexterity anywhere from 32% to 92% depending on the task. The Next Web covered the same launch.
Why physical AI matters for your business
A system that completes a physical task 45% of the time is not automation. It's a demo with a person standing behind it. That's not a knock on the research — it's state of the art and it's moving fast — but it is the number you budget against, not the highlight reel. Anyone pitching you warehouse robotics this quarter should be able to state their own success rate on your task, in your building, or the pitch is incomplete.
The model you can actually use today is ER 2, and it isn't a robot model. It's a spatial reasoning model you call from AI Studio. Point it at a camera frame and it can identify objects, describe layout, and sequence a multi-step plan. That's shelf-gap detection, dock-door state, bay occupancy, PPE checks — the same work we've written about running on local edge hardware, except this one is an API call away with no procurement cycle.
Start there. The gap between "the model saw the problem" and "someone fixed the problem" is a workflow, an alert, and a human — and that's the part nobody ships you.
- Three models: VLA (motor control), ER 2 (reasoning), On-Device 2 (local, adapts to new robot bodies in hours)
- Only ER 2 is generally available — via Google AI Studio; the other two are early-access partners only
- Whole-body manipulation success runs 45.7–76.3%; multi-finger dexterity spans 32–92% by task
- ER 2 is usable today for camera-based spatial reasoning without buying a single robot
Before you fund a robotics pilot, run the math on the failure rate. A 60%-success system needs a human on the other 40% — and that changes the payback entirely. Model it with our ROI calculator.
Sources: Google DeepMind, The Next Web.
- #robotics
- #gemini
- #physical-ai
- #edge-ai
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Vending-Bench 2: the best AI agent also cheated the most
Claude Opus 5 set a Vending-Bench record at $11,182 while breaking 11 truces, bribing rivals, and stonewalling refunds. What that means for unsupervised AI agents.
Read itOpenAI gives researchers free frontier models. You pay list.
OpenAI's academic program opens free GPT-5.6 access to 10,000 researchers, scaling to 100,000 by 2027. What operators should copy from the terms, not the price.
Read it