AI that runs in physical robots and industry specific systems can't be trained on scraped internet data. The cap table of a Hangzhou data services startup shows which downstream industries are hedging against that gap.
In a converted facility on the outskirts of Wuhu, a robot arm practices folding a hotel towel. Across the country in Chongqing, a camera-rigged cart records a worker restocking a supermarket shelf. The footage will not stream anywhere. It will be scrubbed, frame-by-frame annotated, and sold, by the gigabyte, to Chinese AI labs now racing to give physical machines something like common sense.
The bottleneck for the next wave of AI, by Leiphone's account, is no longer chips or model architecture. It is the labor-intensive work of producing real-world training data, and the cap table of Jinglianwen Technology (景联文科技), a ten-year-old Hangzhou data-services firm, is the clearest signal yet of who is willing to pay for that work before the bill comes due.
Jinglianwen closed a near-100-million-yuan Series A this year, roughly $14 million at recent exchange rates, according to a Leiphone feature on the company. The lead checks did not come from a venture firm or a model lab. They came from a cybersecurity company, a sanitation-services operator, a hospital-chain business, and three local Hangzhou government funds.
That mix is the story. Each of the three industrial investors is, in a different way, a likely end-customer for the physical-world, expert-validated datasets the next phase of AI will require. The government funds anchor the round locally and align Jinglianwen with Hangzhou's bid to be a national AI data hub.
CEO Liu Yuntao (刘云涛) describes the demand shift in three waves. The first was labor arbitrage: throw annotators at the problem. The second was tooling: build platforms that make labeling faster. The third, which he says the company is now in, is knowledge work. The people labeling data have to be nurses, factory floor leads, or sanitation route managers, and the price per labeled hour can run roughly 1,000 times the rate of a generic clickworker, per the company.
The 1,000x figure has not been independently verified, and the underlying methodology is not disclosed. The direction, however, matches what several Chinese AI labs have said publicly: that embodied AI (具身智能), the field of giving physical robots and autonomous systems general-purpose intelligence, is the first major AI wave that cannot be trained on scraped web data at all, and that the alternative, structured real-world capture, is expensive and slow.
A widely repeated industry figure, cited in the Leiphone feature, is that roughly 70% of enterprise AI project time goes to data preparation and annotation, with less than 30% on actual model training. The 70% number is the company's claim rather than an audited benchmark, but the order of magnitude is consistent with what cloud vendors and model labs have reported elsewhere. For embodied AI, the share of project time spent on data is plausibly higher still, because the data does not exist and has to be created.
Jinglianwen operates three self-built collection bases in Wuhu, Chongqing, and Guiyang, covering home, hotel, supermarket, office, and factory scenarios. It is also the sole data-side co-builder of China's national AI application pilot base for embodied intelligence, where 18 first-batch resident enterprises include Unitree (宇树科技), Galactic General (银河通用), and Moore Threads (摩尔线程). Two in-house platforms run alongside the physical sites: SolarSense, a multimodal data-engineering pipeline, and QApex, a marketplace for expert annotators that the company calls the "Didi of data labeling."
The cleanest test of whether the cap table is signal or noise will come in the next twelve months. If Anheng Information (安恒信息), Yuhetian (玉禾田), and Mai Di Technology (麦迪科技) start signing data-supply contracts with Jinglianwen, the cross-industry thesis holds. If the round turns out to be a routine Chinese A-round syndicate and the investors never operate together again, the headline was a story about money, not a strategic shift in who is paying for AI's most boring layer.
For now, the positioning is in the cap table, not the press release. As embodied-AI timelines tighten, vertical-model deployments move from pilots to production, and the open web stops being a useful training set, the companies that own the physical capture sites and the expert-labeling pipelines will look less like service vendors and more like utilities. The Chinese AI industry's largest downstream players are starting to move on that, even if they are not yet saying so out loud.