At WRC 2026, China's largest robotics trade show, a robot barista ran a real café for five days, testing a physics prediction system — a latent space world model — originally developed for autonomous driving as a brain for uncontrolled real world
A wheeled robot in a canvas apron picked up a crumpled napkin from a café table, dropped it in a bin, then rolled its gripper under a UV sterilizer. Five days earlier it had poured latte art for a stranger. The loop, napkin to UV, ran continuously at the World Robot Conference in Beijing this August, one of the few booths where the robot was doing real work for real customers rather than running a scripted demo.
The WRC, China's largest annual robotics trade show, drew hundreds of embodied-AI companies this year. Embodied AI is the umbrella term for AI that controls robots moving through and acting on the physical world. Most companies staged tightly choreographed routines. Wujie Power, a Beijing startup founded in March 2025, co-located a working coffee shop with Korean chain Hollys Coffee and let customers walk in off the exhibition floor. The robot had no scripted paths. People moved freely, placed orders randomly, and the machine had to navigate around them while identifying private items on tables it should not touch, a level of unstructured operation that most embodied-AI demos avoid.
The café was a test of one question: can a robot that has never seen a customer's kitchen still wash their dishes? Wujie Power's bet is that the answer comes from the engineering playbook of autonomous driving. The company calls its approach a "latent-space world model." Instead of predicting the next frame of pixels, the system predicts how the physics of the room changes when the robot acts: where the customer will be two seconds from now, which objects will shift, what the gripper will feel. The CEO, Zhang Yufeng, frames this as a direct transplant of work he did at Horizon Robotics, where he led the Journey chip family and ADAS software to the top share in China's independent-brand passenger-car market.
The dominant embodied-AI paradigm today is the VLA, or vision-language-action model, which trains a robot by imitating human demonstrations. Wujie Power's published position is blunt: "VLA is imitation learning. It does not generalize to open environments." A napkin dropped in an unexpected spot, a customer standing in the wrong place, a cup rotated by a stranger, each is a distribution shift the VLA has never seen. The latent-space world model, by contrast, reasons about the physics and recovers.
Others in the industry take a different approach, building pixel-level generative world models that simulate reality frame by frame. Wujie Power's wager is that physics reasoning in a compact latent space is more efficient and more robust to the long tail of café behavior. The WRC demo was the first public test of that bet in an uncontrolled room, and the first batch of K15 robots has already shipped to Europe after the model received EU industrial-grade CE certification in July. The company calls that milestone a first for an embodied-AI robot, though the claim has not been independently verified.
The company says its order book stands at 700 million yuan, roughly $98 million at approximate mid-2026 exchange rates, and that it raised hundreds of millions of dollars in the first half of 2026 from a mix of industrial and financial investors. Both figures are company-sourced and unconfirmed. What is confirmed is that the WRC main forum formally established a World Model Expert Committee under the China Electronics Society this August, with Zhang appointed to its roster, and that the panel's first public debate was titled "World Model to Real-World Distance."
That distance is what the café was built to measure. If the autonomous-driving playbook holds in an open café, embodied AI has just bought several years. If it does not, the napkin-to-UV loop is a one-off stunt, and the industry's real product is still years from the home. The next K15 shipments go to European logistics partners in Q4.