Skip to content
Rush Commerce
AI & Automation3 min read

Gemini Robotics 2: read the success rates, not the demo

Google DeepMind shipped three Gemini Robotics 2 models. Only one is open to developers, and the published success rates are the number that matters.

Google DeepMind released Gemini Robotics 2 today — three models covering whole-body humanoid control, embodied reasoning, and on-device inference. The demo videos are impressive. The success rate tables published alongside them are more useful, and they say something different.

What actually happened

Per DeepMind's announcement, three models shipped together:

Gemini Robotics 2 is the vision-language-action model — it converts camera and language input into motor control, driving full humanoids from feet to fingertips as well as bi-arm setups. Gemini Robotics ER 2 is the embodied reasoning model: it reads a physical environment, plans multi-step tasks running several minutes, and coordinates multiple robots working together. Gemini Robotics On-Device 2 runs locally with no internet connection and adapts to an entirely new robot body with a few hours of data.

Access is tiered, and the tiers matter more than the capabilities. ER 2 is available now in Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. The other two — the VLA and the on-device model — are early-access partners only, behind a sign-up form.

Now the numbers. DeepMind's own reporting puts whole-body manipulation success at 45.7–76.3%, gripper dexterity at 74.2–89.6%, and multi-finger dexterity anywhere from 32% to 92% depending on the task. The Next Web covered the same launch.

Why physical AI matters for your business

A system that completes a physical task 45% of the time is not automation. It's a demo with a person standing behind it. That's not a knock on the research — it's state of the art and it's moving fast — but it is the number you budget against, not the highlight reel. Anyone pitching you warehouse robotics this quarter should be able to state their own success rate on your task, in your building, or the pitch is incomplete.

The model you can actually use today is ER 2, and it isn't a robot model. It's a spatial reasoning model you call from AI Studio. Point it at a camera frame and it can identify objects, describe layout, and sequence a multi-step plan. That's shelf-gap detection, dock-door state, bay occupancy, PPE checks — the same work we've written about running on local edge hardware, except this one is an API call away with no procurement cycle.

Start there. The gap between "the model saw the problem" and "someone fixed the problem" is a workflow, an alert, and a human — and that's the part nobody ships you.

  • Three models: VLA (motor control), ER 2 (reasoning), On-Device 2 (local, adapts to new robot bodies in hours)
  • Only ER 2 is generally available — via Google AI Studio; the other two are early-access partners only
  • Whole-body manipulation success runs 45.7–76.3%; multi-finger dexterity spans 32–92% by task
  • ER 2 is usable today for camera-based spatial reasoning without buying a single robot

Before you fund a robotics pilot, run the math on the failure rate. A 60%-success system needs a human on the other 40% — and that changes the payback entirely. Model it with our ROI calculator.

Sources: Google DeepMind, The Next Web.

  • #google
  • #robotics
  • #gemini
  • #physical-ai
  • #edge-ai
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.