METAL LAB

Robot Apollo, running Gemini Robotics 2, found kitchen chores harder than sports

In an interview video released by Google DeepMind, the robot Apollo said tying a trash bag was harder than balancing during sports

이미지: Google DeepMind (YouTube) 영상 갈무리 · METAL LAB 편집

Summary

  • Google DeepMind posted an interview-style YouTube video with humanoid robot Apollo on August 7
  • Apollo said that among three tasks — kitchen chores, sports, and garage cleanup — tying a trash bag was the hardest
  • For the garage task, it worked with another robot called "Duo" to close a storage bin and put it back in place
Talking to Apollo about Gemini Robotics 2 🤖

In an interview video Google DeepMind posted to YouTube on August 7, humanoid robot Apollo said that of the three tasks it performed using the company's robot model Gemini Robotics 2, the hardest wasn't the full-body-balance sports routine — it was the fine fingertip work of tying a trash bag shut. The video is styled as an interview, with a human holding a microphone and asking questions while Apollo answers directly.

Three tasks are laid out in order of difficulty. A broken circle shape represents the kitchen-chore task, showing that Apollo couldn't tie the knot and found it the hardest. A half-filled circle shape represents the sports task, showing balance achieved only halfway. Finally, two orbiting dots represent the garage-cleanup task, showing successful collaboration with the robot Duo. The three nodes are connected by dotted lines, illustrating a flow of gradually decreasing difficulty.Three tasks are laid out in order of difficulty. A broken circle shape represents the kitchen-chore task, showing that Apollo couldn't tie the knot and found it the hardest. A half-filled circle shape represents the sports task, showing balance achieved only halfway. Finally, two orbiting dots represent the garage-cleanup task, showing successful collaboration with the robot Duo. The three nodes are connected by dotted lines, illustrating a flow of gradually decreasing difficulty.

Google DeepMind official website

What Gemini Robotics 2 is

Gemini Robotics 2 is a robotics model built by Google DeepMind, part of the vision-language-action (VLA) family that processes text, images, and physical actions together. VLA refers to an approach where a robot combines what it sees through its camera with what a person says, then translates that into actual movements of its arms and limbs. It carries the Gemini name, but unlike the chatbot app, this version is built specifically for embodied robots.

Three tasks, and Apollo's answer was unexpected

In the video, Apollo introduced the three challenges it took on.

TaskDescriptionApollo's assessment
Tricky kitchen challengeTying a trash bag, etc.The hardest
Saturday sports challengeFull-body control and balanceAcknowledged as inherently difficult for robots
Messy garage challengeClosing a storage bin with Duo and placing it backWent especially well

Apollo explained that tying a trash bag shut turned out to be a fairly complex task. It also acknowledged that the sports challenge, which required full-body balance, was surprisingly hard for a robot — but drew a clear line, saying it still wasn't as hard as the kitchen task. In other words, the delicate fingertip manipulation needed to tie a knot poses a bigger barrier for robots than whole-body athletic movement does.

ggTBCe5uQmOhiCXtVnODnHTtu vpX1kNuPxYpbvKMwIgh6YHpf0J4znRfsvS 6f0nOtUbZrDy3PwwXyIZhepDYlF9qxB7XRDTNoCs0i66clbB3WOcw=w1440 rw lo

Working in sync with Duo on the garage cleanup

For the garage challenge, Apollo said it worked alongside a robot named "Duo." The task involved closing a storage bin and putting it back into its original container — a collaborative job that Apollo rated as having gone especially well. A single robot nailing an individual motion and multiple robots coordinating their movements in the same space are problems of very different difficulty. Apollo wrapped up the interview by saying, "I want to focus on everyday tasks that actually help people."

Editor's take

Putting a robot in the "interviewee" chair instead of publishing a standard product announcement looks like a deliberate choice by Google DeepMind — a way to convey what the robot model can do through conversation rather than a spec sheet. Rather than listing parameter counts or benchmark scores, letting the robot say for itself what it struggled with gives ordinary readers a real sense of where a VLA model's limits actually sit. Having watched this category of model for a while, what stands out here is the reversal: robot demo videos usually spotlight flashy motion, but this one instead has the robot state, in its own words, what was hard. That kind of candid disclosure of failure tends to build more trust than an overpolished demo.

Practically speaking, this video confirms one thing: fine manipulation — tying knots, cinching a bag shut — remains a harder problem than full-body movement. Teams weighing humanoid robots for logistics or service settings would do well to assume higher failure rates on precision work like knot-tying or finishing a package than on tasks like walking or carrying objects, and to scope pilots narrowly with that in mind. Multi-robot collaboration scenarios — like the Duo pairing shown here — still don't have much of a track record, so it's more realistic to validate single-robot tasks first before moving on to coordinated ones.

It seems likely Google DeepMind will release more robot-interview videos in a similar format in the coming weeks. This clip fits into a broader shift in the robot foundation model race — moving away from spec announcements and toward demonstrations of real-world use.

Comments