Teleoperation datasets are about 100,000 times smaller than the corpora behind today's language models, and hiring more remote operators can't close the gap.
Humanoid robots are pitched as the answer to labor shortages. To train them, the industry is hiring humans to control the hardware remotely, one demonstration at a time. The size of that workforce, more than any single demo, is the clearest signal that the dominant training method is running out of room before it can replace the labor it is built on.
Teleoperation datasets are roughly 100,000 times smaller than the corpora used to train today's language and vision models, according to an analysis in The Robot Report. Frontier language models train on trillions of tokens; a teleoperation corpus is millions of short, expensive, hand-recorded episodes. The gap cannot close by hiring more operators, because the world keeps changing: every new room, object, or lighting condition asks for a fresh demonstration.
A teleoperator wearing a headset cannot feel what the robot is touching, and depth perception through a camera stream is unreliable enough that operators move slowly and overcorrect. The footage a robot learns from is a video of someone struggling with a controller. Scaling that footage doesn't buy the dexterity the original humanoid pitch promised.
A commercial teleop-data ecosystem has emerged across China, India, Europe, and the US, with startups selling recorded demonstrations the way companies once sold labeled text for early language models. The Wired profile of Flexion and the Humanoids Daily funding write-up sit inside this same labor picture, even when they are framed around a single company's technology. The operator pool's geography, not any single lab's GPU count, is the new leading indicator for how the industry is scaling, and the map of that pool looks more like the early-labeling markets for language models than the map of where humanoid labs are headquartered.
The original humanoid pitch has been demographic: aging populations and shrinking workforces will leave jobs that humans cannot fill. The training data for those robots, today, is being produced by humans, at wages that vary by geography, in workflows designed to mimic the very jobs the robots are meant to replace. That is a measurement of how far the dominant method is from substituting for the labor it depends on, not a moral sidebar.
Flexion, named in the Robot Report piece as an exemplar of one such alternative, is building a reinforcement learning and sim-to-real platform: training policies inside simulated environments and transferring them to hardware, the way a fighter pilot learns in a simulator before a real cockpit. The company has raised $50 million to pursue that approach, and its Reflect v1.0 release notes describe a system designed to reduce the amount of human-recorded footage a robot needs. The $50 million figure traces to Humanoids Daily and hasn't been confirmed in a direct company filing; the more useful question is what any alternative path has to beat.
Three things would break the wall thesis: a leap in sim-to-real transfer that lets policies trained in simulation work on hardware without re-demonstration, a breakthrough in unsupervised world models that learns from passive video, or a synthetic-data pipeline that generates useful training signal without humans in the loop. Until any of those arrives, the next 18 months of humanoid news will read two ways at once. A new training milestone could be progress toward generalizable autonomy, or it could be the dominant approach being patched rather than solved. The teleop labor footprint, not the demo video, is how to tell which.