Google DeepMind's Gemini Robotics 2 model runs the same control software on Apptronik's 5'8" bipedal Apollo 2 humanoid and Franka's tabletop F3 Duo dual arm rig.
A Google-built AI model is folding laundry, sorting trash, and clearing dishes on a 5'8" humanoid in a YouTube clip released this week, and the same software is threading a needle on a tabletop arm rig in Germany. Both runs use the same control software. That combination is the architectural bet underneath the chore clip.
Google DeepMind launched Gemini Robotics 2 on Monday as a "whole-body" control layer for robots, covering motion from feet to fingertips, plus fine-grained dexterity and multi-robot coordination. A companion model, Gemini Robotics-ER 2, runs above it as the planner: it chats with humans, breaks down a task like "tidy the kitchen" into steps, and hands motor execution to the lower-level model. The ER variant can also call external tools, including Google Search, mid-task.
The launch included two simultaneous demonstrations on two structurally different bodies: Apptronik's Apollo 2 humanoid, a 5'8" bipedal platform offered in both bipedal and wheeled-base configurations, and Franka's F3 Duo, a stationary dual-arm rig used in academic and industrial research. Same model, two embodiments. The pairing is the architectural claim the company is asking readers to evaluate.
The test is whether a single set of learned control weights can transfer across robots with different kinematics, sensors, and physical reach. A chore video is a clip. A chore video plus a separate demonstration on a structurally different body, using the same weights, is closer to a generalization claim. The released videos show the model succeeding on both, in scripted conditions. The question the demos do not answer is whether the same weights would hold up on either body in a new environment, with a novel object, in front of a non-engineer.
Apptronik is feeding that test. Additional Robot Park sites are planned at customer and partner locations. The loop is direct: data from real Apollo 2 deployments flows back into the model, which Apptronik then deploys in the next generation of the same hardware. That tightens the path between training and product, but it also narrows the diversity of embodiments the model sees outside the lab.
The chore clips are staged, the lighting is controlled, and the failure modes that matter in a real home (a wrinkled shirt, a partially opened jar, a pet underfoot) are not the failure modes the videos show. No independent benchmark for in-home success rates has been published, and the safety claims in the announcement are company-asserted. The model's ability to plan multi-step tasks through the ER 2 layer is a research claim, not a deployment claim; the announcement frames both as "near future" capabilities, with the model card and full technical details still to be read in full.
A cross-embodiment evaluation on Apollo 2 and F3 Duo against a held-out task set, with results that include failure cases and a clear metric, would be the next checkpoint. Robot Park's data loop, if its output is shared with outside researchers, would be a second. The chore video is the surface. The architecture is the claim.