Yuanli Lingji's open DM0.5 topped RoboDojo and held its own on four other robotics control benchmarks, a rare generalist showing in a per scenario tuned field.
A new open-source model for controlling robot arms has posted the strongest average score yet on a multi-institution robotics benchmark, and did it while generalizing across very different jobs: remembering a sequence of covered blocks, sorting a chaotic conveyor belt, and parsing a short human demonstration video into a long action sequence. The interesting question is not that one team topped one leaderboard, but whether a single generalist foundation model plus a brief human demo is finally enough to retire the per-scenario tuning that has defined industrial robotics for a decade.
DM0.5 is a "robot brain": a foundation model for controlling robot arms and dexterous hands, not a chatbot. It comes from Yuanli Lingji (原力灵机), a Chinese embodied-AI lab, and was reported this month by QbitAI as having taken the top spot on RoboDojo, a robotics control benchmark jointly run by the University of Hong Kong's Multimedia Lab, UC Berkeley, Tsinghua, and roughly twenty other academic groups. RoboDojo is the kind of leaderboard the field treats as a stress test: as of July 2026, the top simulation models on the board averaged 8.80% task success, and the real-machine leaderboard sat at 12.8%. A 19.34% average for DM0.5 sits well above the previous ceiling on simulation tasks.
The generalist part is what the vendor is betting on. A single DM0.5 base model is reported to handle six scenario classes at once: learning from human teaching, fine manipulation, long-horizon logistics, anti-interference, flexible-object handling, and long-memory tasks. That is unusual in a field where most production systems are still engineered one task at a time. At a World Robot Conference demonstration in Beijing's Yizhuang district, a DM0.5-controlled robot sorted packages at roughly three seconds per package with above 99% single-piece separation accuracy on a real conveyor belt, and the team said no moves were pre-scripted. A separate fine-manipulation showcase built a 3.5-meter model Great Wall from 81,920 mini-blocks across fifteen hours using six robots, with grasping and snapping precision of 0.1 to 1 millimeter, below the 0.3 to 1 millimeter tremor band of a natural human hand.
DM0.5 also leads the RoboDojo Memory dimension by 47.74 points, and the QbitAI writeup notes a 100% success rate on the Cover Blocks memory task across three random trials. The model carries a native 60-second memory window, which the vendor says is what lets a robot watch a human demonstration video and chain the resulting action sequence into a long-horizon task. Under the hood, DM0.5 splits cognition into a slow System2 path for high-level reasoning, refusal, and zero-shot prediction, and a fast System1 action-expert network for reflexes inside long-horizon work. The dual architecture is presented as the anti-disturbance mechanism that keeps the system from drifting when something unexpected happens mid-task.
Cross-benchmark numbers are equally the story. Per the same QbitAI report, DM0.5 posts 99.0% overall success on LIBERO, 93.6% and 93.3% on the simple and complex tracks of RoboTwin 2.0, 89.0% / 53.6% / 44.1% across the three difficulty tiers of VLA-Arena, and a 54.42 score with 43% success on RoboChallenge Table30 v2. Each of those leaderboards measures a different slice of embodied competence, and holding rank across all of them is the kind of result that does not usually happen with one base model.
That is the constructive case. The honest hedge sits in the same source. The 19.34% average is strong relative to a leaderboard where the previous top simulation models averaged 8.80%, but it is still very low for any industrial deployment that does not include human-in-the-loop and per-task error correction. QbitAI itself frames the field as still needing a "ChatGPT moment for robotics," and the same coverage preserves the lab's own constructive skepticism about data-scaling orthodoxy: if a single base model can generalize across very different jobs, then the field's "millions of hours of robot data" race may be answering the wrong first-principles question. The vendor is open-sourcing the weights and code so that claim can be audited and replicated. Until a third-party lab reproduces the cross-benchmark numbers and pushes the average success rate toward something a factory floor can tolerate, the generalist story is the most interesting read on this release, not the rank itself.