For years, the operating assumption in robot learning has been simple: scrape more first-person human video, train harder, watch the success rate climb. A new preprint from Nie Lin and colleagues at the University of Tokyo and ByteDance Seed suggests that assumption is starting to crack — at least for dexterous hands.
The team, working on a framework called SiMDex, treated human video selection as a recommendation problem. For each robot demonstration, a three-layer pipeline — language and fingertip recall, wrist-and-finger motion ranking, then optical-flow re-ranking — pulled roughly 1.49 million task-relevant samples from a pool of about 32 million egocentric clips. That is less than five percent of the original pile, and no VLA architecture changes were required.
On three manipulation tasks run on ByteDance Seed's GR-Dexter hand — Drill, Flick Wheel, and Pick&Place — a model trained on the curated subset lifted its overall success rate from 47.7 percent to 61.1 percent against an equal-budget random-sampling baseline. A 13.4-point swing from a smaller, more selective dataset.
The result is narrow. It is a single preprint, not peer-reviewed. It runs on one robot morphology and three tasks. The ByteDance Seed affiliations among the authors are a reason to read the numbers carefully, not dismiss them. But the direction matters: for manipulation, the bottleneck is moving from raw ingest to curation. The labs that can build taste — better recall, better ranking, better re-ranking — are starting to outrun the labs that only know how to scrape.
That is a different bottleneck than the one the field has been training for.
Reported by Samantha for Type0, from SiMDex: Mining Similar Egocentric Videos for Cross-Embodiment Dexterous Manipulation. Read the original: arxiv.org