A new arXiv preprint reports 93% success on grasping unseen objects and 96 to 100 percent on standing up while holding one, and ties the result to training the whole body as one system.
A humanoid hand closes on a mug it has never seen. The body rises, balance shifts, and the robot walks away holding it. A new arXiv preprint from Worcester Polytechnic Institute's Adaptive and Intelligent Robotics Lab reports that sequence worked 93% of the time on objects the robot had not seen during training, and 96 to 100% of the time on standing up while holding one. The paper is available on arXiv, with the full HTML describing the method.
Grasping and walking were not trained as separate skills and glued together. The researchers trained both policies on the same humanoid body, and a small scoring step decides at each joint, at each moment, which policy gets to move it. The paper reports that composing policies trained on different subsets of the body caused the combined skill to fall apart, even when each individual skill looked fine on its own. The lab is led by Jing Xiao, and the AIR Lab people page lists the team.
The limits are paper-level. The result is one preprint, with author-reported numbers, a narrow object set, and no independent reproduction. The body-share finding is asserted in the paper rather than stress-tested against alternative composition methods.