TriWorldBench, built by Peking, Tsinghua, Beihang, Shanghai Jiao Tong, and the University of Science and Technology of China (USTC), tests whether a robot's head camera and two wrist cameras can tell a coherent story about the same scene.
Peking University, Tsinghua, Beihang, Shanghai Jiao Tong, and the University of Science and Technology of China jointly published the first weekly ranking for TriWorldBench, a benchmark that judges robot "world models" by whether they understand a scene, not just render a believable video. The leaderboard lives at triworldbench.com.
A world model is the AI inside a robot that tries to imagine and predict its physical surroundings from camera input. TriWorldBench focuses on three of those cameras: a head view plus left- and right-wrist views on a manipulation robot, and checks whether the three streams agree on what the robot is doing, what it is grasping, and what the target object looks like.
The benchmark covers 500 synchronized three-view episodes across 50 robot manipulation tasks, scored across six dimensions (multi-view consistency, task alignment, physics and 3D consistency, motion quality, temporal consistency, and visual quality) and aggregated into a single TWB-Score from 19 underlying signals. The shift matters because most prior evaluation rewarded visual realism over physical and cross-view consistency.
Within a week of launch the leaderboard drew 14 official teams, more than 10,000 page views, and over 200 evaluation-tool downloads. Rankings update continuously.
The second weekly ranking will show whether the gap between video generation and world understanding is closing or widening.