EgoVerse, Danfei Xu's Georgia Tech coalition, has 1,362 hours of first person video and 1,965 tasks, betting shared data can replace the remote controlled robot demonstrations each lab currently collects, with Lightwheel, a Beijing physical AI
Robot training depends on teleoperated demonstrations, where a human operator drives a robot arm through a task. Humans generate the same kind of visual-motor-task knowledge for free, every time they cook a meal, fold laundry, or unscrew a battery. A new international coalition called EgoVerse is betting the second stream can be turned into the first, and that doing it as a shared standard, the way ImageNet was for vision models, will be the substrate that finally lets robot policies generalize across labs and hardware.
Hand footage captured on Meta's Project Aria glasses or a smartphone can be re-played on a robot arm without per-body retraining, in a transfer mode the project calls cross-embodiment. The claim collapses if those policies fail to transfer to a humanoid arm outside the training distribution. If the mechanism holds, EgoVerse is the missing substrate. If it does not, embodied AI stays stuck in single-lab demos.
Danfei Xu, a former Stanford PhD student co-advised by Fei-Fei Li and Silvio Savarese, now leads the Robot Learning and Reasoning Lab (RL2) at Georgia Tech and is the project's principal organizer. His coalition's most recent update, reported this month by Chinese tech outlet QbitAI, puts the dataset at 1,362 hours of first-person human demonstration, roughly 80,000 episodes spanning 1,965 tasks across 240 scenes, collected by 2,087 contributors. The project paper describes EgoVerse as a "first-person human operation data ecosystem," the kind of shared infrastructure ImageNet was for labeled photographs.
Stanford's Universal Manipulation Interface (UMI), built by Shuran Song's Robotics and Embodied AI Lab (REAL), is a handheld gripper that records the same kinematics a robot would execute. When EgoVerse footage is replayed through a UMI-trained policy, the resulting behavior transfers, in the project's claim, to a stationary or mobile manipulator without per-body fine-tuning. Georgia Tech validates object-in-container tasks on this loop; Stanford focuses on bi-manual fine manipulation, the two-handed kind like threading a cable or opening a jar; UC San Diego and ETH Zurich extend the protocol to long-horizon bi-manual tasks, the kind that chain ten or more steps without resetting.
Lightwheel, a Beijing-based physical-AI company, is the only Chinese industry member in the alliance. Lightwheel is not a humanoid-robot OEM, the category most Chinese embodied-AI coverage tracks; the company's most recent public demonstration is a physical-AI education system shown at WAIC 2026 in Shanghai. A working EgoVerse would let Lightwheel train policies without owning a robot fleet. The Chinese seat on a US-led academic coalition is itself a watch-list item as the standards-setting window closes.
Meta's Project Aria supplies purpose-built first-person capture: an RGB camera, simultaneous localization and mapping (the technique that lets a device track its own motion through space), and hand-tracking sensors. Mecka AI pushes the opposite end of the spectrum, smartphone video with no special rig, which is the only way the dataset reaches hobbyist and household contributors. Scale AI handles annotation. A policy that transfers across Aria, smartphone, and UMI captures would not be tied to one sensor stack, which is why the coalition runs all three collection pipelines in parallel.
Linux Foundation's Newton project is a separate open-source robot learning initiative focused on standardizing simulation and runtime layers, not the data layer. EgoVerse and Newton do not conflict, but they also do not yet reinforce each other. A cross-institution data coalition can drift toward single-vendor capture if the partners whose hardware and annotation pipelines do the most work end up setting the spec. The consortium's current academic-organizer structure, with Georgia Tech as the backbone and Stanford, ETH Zurich, and UC San Diego handling validation, will be tested as new partners join.
Whether the substrate works will be settled when EgoVerse-trained policies run on a humanoid platform outside the project's own lab distribution. The coalition has signaled an updated scale release and broader partner validation tasks in the coming quarters.