Embodied AI, the software that lets robots perceive and act in physical space, is claiming top spots on public benchmarks over Nvidia and Ant Group; the harder fight is teaching it to work in unfamiliar homes.
Embodied AI, the software that lets robots perceive and act in physical space, is posting benchmark wins over Nvidia and Ant Group; the harder fight is teaching it to work in unfamiliar homes.
China's embodied-AI startups are winning the benchmark race and raising fresh capital. The hard part, getting robots to operate in unfamiliar physical environments, is still waiting on a training-data flywheel that neither of the country's two most-watched founders has cracked. The same chairman making the loudest forecast is the one conceding broad deployment is still 4–5 years past it.
ACE Robotics chairman Wang Xiaogang, who also co-founded the Chinese AI company SenseTime, told Reuters he expects embodied intelligence to reach a "ChatGPT moment" by the end of 2027, driven by world models (AI systems that simulate physical environments to help robots plan action) and large-scale environmental data capture. The same chairman conceded that even a 2027 inflection point would still leave 4–5 years before broad commercial implementation across sectors. The startup, founded in July 2025, has raised more than $100 million in the first half of 2026 across several financing rounds, with backers including Ant Group, the Chinese fintech affiliate of Alibaba, and SenseTime.
On paper, ACE's Kairos-4B open-source model, a 4-billion-parameter system hosted publicly on GitHub and Hugging Face, ranks first on public embodied-intelligence benchmarks according to the company, outperforming Nvidia's Cosmos 3 and Ant Group's Lingbot. The model integrates perception, multi-modal understanding, physical simulation, action planning, and long-horizon video and action prediction. The ranking is a self-published claim pending independent replication on the same tests.
Wang Xingxing has separately told Reuters he expects a robot-brain breakthrough in 2–3 years at the earliest. Both founders land on the same bottleneck: high-quality, real-life training data.
Public benchmarks are a contested measurement. They reward models that fit a fixed test distribution, not robots that have to pick up a stranger's coffee cup in a stranger's kitchen. The work of capturing messy, real-world physical interaction (opening a sticky drawer, navigating a cluttered hallway, recovering from a dropped object) does not scale by scraping the web. It requires teleoperated demonstration, which is expensive and slow, instrumented physical environments, which do not exist at the scale needed, or synthetic data grounded in physics. Simulators that produce useful training data are themselves a hard research problem, because small errors in physical realism compound into policy failures when a robot transfers to the real world. Chinese labs have not yet shown a flywheel that produces such data at industrial volume.
Investor scrutiny in China's humanoid sector is shifting accordingly. The early story was demo theater: choreographed dance routines, athletic feats, viral video clips. As valuations come under review, the question that matters is whether a robot can perform useful work in an unfamiliar setting, and how much it costs in data, hardware, and engineering to teach it to do so. ACE Robotics has stated an ambition to list "as early as permitted," a window that Chinese listing rules typically tie to three fiscal years of operating history.
The 2027 date is best read as a forecast, not a milestone readers should plan around. Chip supply and model architecture are the parts of the embodied-AI race that attract capital and headlines; training data is the part that resists both, since it does not scale by scraping the web and does not come from a foundry. The falsifier is concrete: if a Chinese lab cracks a synthetic- or teleop-data flywheel at scale, a system that produces enough high-quality real-world training video to teach robots generalizable skills, the line becomes credible. Until then, the data gap, not chip supply or model architecture, is the binding constraint that every "ChatGPT moment" headline assumes away.