At Sydney's RSS 2026 (the Robotics: Science and Systems conference), researchers said embodied AI (AI that controls robots) has no foundational breakthrough anywhere; the 'Chinese version of X' brand, they said, is a hype cycle tell, not a
At Sydney's RSS 2026 this month, embodied AI, the work of putting AI in control of physical robots, heard a sharper diagnosis from the people closest to it. Two doctoral students from Physical Intelligence walked a QbitAI reporter through the same claim in separate conversations: vision-language-action (VLA) models, which map camera images and instructions directly to robot actions, and world models, which learn a simulated environment that an agent can roll forward in its head, are not rivals. They are complements, and the pi0.7 release on arXiv (RSS paper 74) already embeds a world-model component inside a VLA stack. The dichotomy that has dominated Chinese embodied-AI discourse for months, "VLA已死,世界模型当立" ("VLA is dead, world models must rise"), is asking the wrong question.
The provocations came from the PI orbit. A respondent, paraphrased to QbitAI from the Sergey Levine group, walked the field's own argument back onto the conference floor: that the leading North American embodied-AI companies, Physical Intelligence included, are still some distance from mature; that the field has not yet produced a truly foundational piece of work; that China's real lead is in robot hardware; and that if the US has no mature answer, there is no reason for Chinese companies to rush to brand themselves as "the Chinese version of X." A PI intern, asked whether world-model claims to predict action consequences had panned out, cut in: "Have you actually seen it do that?"
The program shape runs against the "VLA is dead" claim in a way that is hard to miss once you look. Monday's marquee slot was a 45-minute joint session on world models and memory, shared because the committee could not separate the two. Tuesday afternoon and evening, from roughly 3 p.m. through the close, were dominated by imitation-learning and VLA reports, the latter half of the day given over to the architecture the Chinese narrative has declared obsolete. Out of 708 submissions, eight papers were finalists for the three best-paper awards, and program chair Matei Ciocarlie's opening keynote was built around a single keyword, "connection": in-person exchange beyond what papers and arXiv/X posts can carry.
What is actually load-bearing, the on-record researchers argued, is data curation, not architecture. Karen Liu's closing keynote framed the field as "Model-rich, Data-poor": accumulated simulation, planning, and vision models are not directly robot policies, but they can be combined as teachers that keep producing training data. Simar Kareer, a core author of EgoVerse, put it more sharply: the gap is not "more data" but "data curation, mixing ratios, and organization." Wenzhen Yuan, on tactile sensing, added that tactile is still mostly bolted on as an additional modality to VLA, with the harder problems being sensor-specific transfer (different tactile sensors do not transfer to each other) and whether the end-effector or fingertip form factor even leaves room for tactile to be useful. The pattern in all three talks is the same: the next bottleneck is integration, not which model class wins.
On the Chinese side of the floor, the story was complementary. Booths from 星海图, 自变量, 千寻, 智元, 宇树, 高擎, and Seeed drew mostly university and research-institute visitors; the first questions were price, ROI, secondary-development conditions, and whether the product could deploy into local scenarios such as mining or inspection. QbitAI's on-site read is blunt: Chinese robot hardware export is clearly ahead of Chinese software and data services at this stage. Hardware is where China has a real wedge; software is where it does not yet have a story of its own, which is part of why the "Chinese version of X" branding pattern is so tempting and so misleading.
The branding pattern is also a tell about cycle lag. RSS 2026 closed on July 13 in Sydney; submissions closed January 30, acceptances went out April 27, a roughly 5.5-month arc from paper to podium. The "VLA is dead, world models reign" line that has circulated through Chinese embodied-AI funding decks for several rounds is closer in cadence to a Twitter discourse than to a top conference. A half-year gap between conference research cadence and industry hype cycles is not unusual. It just is not usually this visible, and it is not usually a country's tech press announcing a verdict on the wrong architecture six months early.
The watch item for the next cycle is not whether VLA or world models "wins." It is whether Chinese teams stop importing a US brand-name slot and start naming the hardware-led, data-curation-heavy work they are actually positioned to do. The original is not finished, and the room in which it might be finished is, for the moment, the larger one.