• Latest
  • Trending
  • All
Google DeepMind announcement image for Gemini Robotics 2 showing an Apollo humanoid robot crouching to grasp a green watering can rendered with a point cloud overlay, beside the headline Gemini Robotics 2 brings whole body intelligence to robots.

Gemini Robotics 2: what Google’s own benchmarks say

31 July 2026
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Thursday, September 17, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Gemini Robotics 2: what Google’s own benchmarks say

by stephane
31 July 2026
in Dev
0
Google DeepMind announcement image for Gemini Robotics 2 showing an Apollo humanoid robot crouching to grasp a green watering can rendered with a point cloud overlay, beside the headline Gemini Robotics 2 brings whole body intelligence to robots.
495
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

The clip everyone shared shows a humanoid crouching to pick a watering can off the floor, walking it across the room, and setting it on a low shelf. Smooth. Slightly unsettling. What almost nobody scrolled to is the chart further down Google's own post, where picking an object off the floor is logged at 45.7 percent. That's the launch in one sentence: DeepMind announced Gemini Robotics 2 on 30 July 2026, three models covering whole-body control, planning and on-device inference, and it published the unflattering success rates right next to the highlight reel. We think that's the most interesting thing about the release.

The short answer

Google DeepMind announced Gemini Robotics 2 on 30 July 2026: a vision-language-action model with whole-body humanoid control, an embodied reasoning model called ER 2 for planning and multi-robot work, and an on-device variant. Only ER 2 is reachable today, through Google AI Studio. The rest sits behind a trusted tester form. The published per-task success rates are the honest part of the launch, and they are lower than the video suggests.

3models in the family
45.7%picking an object off the floor
32%sweeping with a dustpan
Google DeepMind announcement image for Gemini Robotics 2: an Apollo humanoid robot crouching to grasp a green watering can shown with a point cloud overlay, next to the headline Gemini Robotics 2 brings whole body intelligence to robots. Image: Google DeepMind

Three models, and only one you can call

The naming is doing some work here, so let’s separate it out.

Gemini Robotics 2 is the vision-language-action model. Camera in, language in, motor commands out. The headline change from the previous generation is that it drives a whole body rather than a pair of arms bolted to a table. DeepMind’s phrasing is feet to fingertips, and the watering can demo is the illustration: the robot crouches, grasps, stands, walks, places.

Gemini Robotics ER 2 is the embodied reasoning layer. It plans the multi-step version of the task, describes what it’s doing to a person, and this generation coordinates more than one robot on the same job.

Gemini Robotics On-Device 2 is the small one, meant to run locally on the machine instead of round-tripping to a datacentre. DeepMind says it adapts to a new bi-arm embodiment in a few hours, with fewer than 200 examples. Honestly that number impressed us more than any of the demos.

Now the part that changes what you can do this week. ER 2 is in Google AI Studio, plus a private preview on the Gemini Enterprise Agent Platform. The VLA is private preview with a waitlist form. On-Device 2 is trusted testers. So the piece you can put your hands on is the planner, and the piece that actually moves a joint is a form you fill in and wait.

Answer card summarising the Gemini Robotics 2 launch of 30 July 2026: three models covering whole-body vision-language-action control, embodied reasoning with ER 2, and on-device inference, with published success rates of 45.7 percent for picking an object off the floor and 32 percent for sweeping with a dustpan, and only one of the three models callable today.
The announcement, and the two numbers from the same post that the announcement did not lead with.

Read the table, not the reel

DeepMind published per-task success rates. Give them credit for that, because plenty of robotics launches ship a video and nothing else.

On an Apptronik Apollo with Inspire hands, doing whole-body manipulation: 76.3 percent picking from a shelf, 68.4 from a table, 45.7 from the floor. That ordering makes physical sense. A shelf puts the object near chest height with the body already stable. The floor means crouching, which moves the centre of mass while you’re trying to close a hand precisely, and the score drops by thirty points.

Then the fine motor set, Apollo with SharpaWave hands. Unscrewing a light bulb: 92 percent. Screwing one in: 36.

Same bulb. Same hands. Same model.

That pair is the most informative thing in the whole release, and I don’t think it got enough attention. Taking a threaded thing out is a coarse, forgiving motion. Putting it back requires aligning an axis you can’t see well, then applying steady torque without cross-threading, and any wobble in the standing body propagates all the way to the fingertips. Assembly is harder than disassembly, which is a very old lesson in robotics and apparently still true when the controller is a frontier model.

The rest of that column: trash bag 44 percent, ziplock 40, dustpan 32.

Bar chart of Gemini Robotics 2 multi-finger dexterity success rates published by Google DeepMind for the Apollo humanoid with SharpaWave hands: unscrewing a bulb 92 percent, tying a trash bag 44 percent, closing a ziplock 40 percent, screwing a bulb in 36 percent and sweeping with a dustpan 32 percent.
Google's numbers, not ours. The bulb pair at the top and bottom is the same object in two directions.

The gripper results on the Franka Duo arms sit much higher: 89.6 percent on precise insertion, 78.9 on diverse tool kitting, 74.2 on general pick and place. So this isn’t a model that misunderstands the goal. Give it a rigid two-finger gripper on a fixed base and it inserts things nine times out of ten. The losses show up when you add a standing body and twenty-odd finger joints to coordinate at once.

Which is roughly what “whole body intelligence” costs, at this point in the curve.

What the demo reel is and isn’t showing

Engadget’s write-up carries the caveat that matters most, and it came from the briefing rather than the blog: the model was specifically trained to perform every task in that video. These aren’t general-purpose machines improvising in a kitchen they’ve never seen.

Watch the grid clip with that in mind and a second detail shows up. Google labels the playback speed on every panel. Most read “Autonomous 1x”, a couple say 1.5x, and one says 4x. Labelling the speed-up is the honest choice, and vendors don’t always make it. It also tells you these robots are slower than a person at the same task, which the raw footage would otherwise hide.

Still frame from the Google DeepMind Gemini Robotics 2 demo video showing six panels of robots working autonomously: a humanoid tidying a bookshelf, a Trossen arm moving a chess piece, an arm ticking an I am a robot checkbox on a screen, an arm stacking crockery on a rack, two arms handling objects on a wall unit, and a coloured arm sorting small parts into a tray, each labelled with its playback speed. Image: Google DeepMind. Watch the full demo on DeepMind’s own video, and check the speed badge on each panel.

One panel has an arm reaching over to tick an “I am a robot” checkbox on a monitor. Cute. Also a decent joke about where the CAPTCHA arms race has landed.

The safety material, and how much to weight it

DeepMind introduced a benchmark called ASIMOV-Agentic with this release. The previous ASIMOV work scored whether a model would refuse a harmful physical instruction. The agentic version targets orchestration: a model that plans, delegates and coordinates has more room to arrive somewhere harmful through a sequence of individually reasonable steps.

DeepMind also says ER 2 is its best model yet at following safety constraints, that it recognises when people are nearby, and that it can call a safety tool to bring a robot to a stop when someone gets too close. Carolina Parada, who heads robotics at DeepMind, framed the current phase as needing to understand the safety question more deeply, which reads to us like an honest hedge rather than a claim of solved.

Worth keeping the frame straight though. This is a vendor reporting its own scores on a benchmark it wrote. That’s normal and it’s still useful, because a published benchmark is something outsiders can eventually run. It just isn’t independent verification, and nothing here has been through one.

Two-column checklist splitting the Gemini Robotics 2 launch into what is reachable today and what is gated: ER 2 open in Google AI Studio, published per-task success rates and the ASIMOV-Agentic benchmark on one side, against the gated VLA and On-Device models, the caveat that the model was trained on the tasks shown, and the unconfirmed report that ER 2 is built on Gemini 3.5 Flash.
Sorting the launch into things you can check, things you can only queue for, and one claim with a single source.

One loose thread: at least one write-up states that ER 2 is built on Gemini 3.5 Flash. DeepMind’s own post doesn’t say that, and the model page doesn’t either. It’s plausible, it fits the latency story, and we’re treating it as unconfirmed until Google writes it down.

So what does this change

If you build robots, not much this week, unless your waitlist form gets picked up. The reachable piece is ER 2 for planning and spatial reasoning, which is genuinely worth an afternoon in AI Studio if you’re doing anything with cameras and physical space.

If you’re trying to judge where humanoids actually are, this release is more useful than most, precisely because the numbers are in it. A machine that puts a bulb in a socket a third of the time is not ready to do that job. A machine that inserts parts with a gripper 89.6 percent of the time is close to being ready for a narrow one. Both facts came from the same announcement on the same day, and only one of them made the headlines.

We track this kind of gap a lot. Our read on the Tesla Optimus V3 reveal went the same way, and the FCC rule that quietly covers robot vacuums is the regulatory half of the same story. If a specific benchmark claim ever needs checking against the source, our tools won’t help you there, but reading the vendor’s own footnotes usually will.

Sources: Google DeepMind, Gemini Robotics 2 brings whole body intelligence to robots, the Gemini Robotics 2 model page, Engadget and MarkTechPost, 30 July 2026. Every success rate quoted here comes from DeepMind’s own published tables. The caveat that the model was trained on the tasks shown in the video, and the Carolina Parada quote, are from Engadget’s briefing. The claim that ER 2 is built on Gemini 3.5 Flash appears in MarkTechPost only and is not stated by Google, so we have marked it unconfirmed. The hero image is DeepMind’s own announcement graphic and the six-panel still is a frame from DeepMind’s demo video, both credited above. The demo video itself is hosted by Google.

Frequently asked questions

What is Gemini Robotics 2?

A family of three models Google DeepMind announced on 30 July 2026. Gemini Robotics 2 is the vision-language-action model that turns camera input and a spoken instruction into motor commands, including whole-body control of a humanoid rather than just arms. Gemini Robotics ER 2 is an embodied reasoning model that plans multi-step tasks, talks back to humans and coordinates more than one robot. Gemini Robotics On-Device 2 is a smaller VLA meant to run locally on the machine.

Can I use Gemini Robotics 2 today?

Only partly. ER 2, the reasoning model, is reachable in Google AI Studio, with a private preview on the Gemini Enterprise Agent Platform. The VLA that actually drives the robot and the on-device version are limited to early-access partners through a trusted tester programme, with a waitlist form. So you can experiment with the planning brain and not with the part that moves anything.

What success rates did DeepMind publish?

On an Apptronik Apollo with Inspire hands doing whole-body manipulation: 76.3 percent picking from a shelf, 68.4 percent from a table, 45.7 percent from the floor. On Apollo with SharpaWave hands doing fine work: 92 percent unscrewing a bulb, 44 percent tying a trash bag, 40 percent closing a ziplock, 36 percent screwing a bulb in, 32 percent sweeping with a dustpan. On Franka Duo arms with grippers: 89.6 percent precise insertion, 78.9 percent tool kitting, 74.2 percent general pick and place.

Are these robots general-purpose now?

No, and Engadget made the point plainly in its write-up: the model was specifically trained to perform every task shown in the video. Adapting to a genuinely new task or a new robot body is a separate step. DeepMind states that On-Device 2 can adapt to a new bi-arm embodiment in a few hours using fewer than 200 examples, which is a real and useful number, but it is still fine-tuning rather than zero-shot competence.

Which robots does it run on?

The published results cover the Apptronik Apollo humanoid with two different hand sets, Inspire and SharpaWave, and the Franka Duo dual-arm setup with a Robotiq gripper. DeepMind also lists Dexmate, SO101 and Trossen platforms, and coverage of the launch mentions Boston Dynamics Spot. The claim is that the same model scales from a desktop arm to a full humanoid.

What is ASIMOV-Agentic?

A safety benchmark DeepMind introduced with this release, aimed at agentic orchestration rather than single actions: whether a model refuses harmful instructions when it is planning and delegating across steps and across robots. DeepMind also says ER 2 is its strongest model so far at following safety constraints, detecting people nearby, and calling a safety tool that brings the robot to a stop. Those are vendor claims on a vendor benchmark, so treat them accordingly.

Tags: aibenchmarksgooglehumanoidnewsrobotics
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that on 15 September 2026 AWS said it is unable to restore access to resources and data hosted exclusively in the Middle East Bahrain region me-south-1 and in the mec1-az2 zone of the UAE region, because the damage spanned multiple Availability Zones and exceeded what multi-AZ services are designed to withstand.

AWS can’t restore me-south-1, six months after the drone strikes

17 September 2026
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.