Humanoid startups borrowed self driving's data flywheel, where each robot's work trains a better model that lets the next do more, but the four saturable tasks they actually demo (pick and place, package reorientation, box packing, clothes folding)
Tesla has logged millions of autonomous miles and still puts safety drivers in the front seat of its Austin robotaxis. The "data flywheel" humanoid and manipulation startups now borrow from self-driving is a structurally different bet, because almost no robotics task outside self-driving is valuable enough to make endless in-domain collection compound. So the demos cluster on the same four saturable tasks: pick-and-place, package reorientation, box packing, and clothes folding.
The pitch behind the borrowing is familiar. Each robot at work generates new data, new data trains better models, better models let the robot do more work, and the loop compounds. The It Can Think! essay argues most robotics deployments don't actually scale that way. They require specialized support teams, and the data they produce is less valuable than the pitch implies. Only self-driving-style problems, where a single task is so valuable that endless in-domain data justifies endless collection, really scale. Everything else is closer to a novelty pump than a flywheel.
Vedder.io sharpens the same point: scale only compounds if it captures novelty, and novelty comes from new tasks and new states, not more robots repeating the same task. Running a novelty pump requires ongoing operational effort and does not scale.
Real-world robotics demands success rates in the 97% to 99.9% plus range to do useful work, and that bar constrains which tasks the industry can attack. When the constraint is that tight, the field naturally clusters around a small set of well-defined jobs. Pick-and-place is the canonical one: a rigid or semi-rigid part, a known target location, a fixture that absorbs small errors. Package reorientation and box packing are variations on the same theme, with a single object type, a small set of orientations, and a fixed destination. Clothes folding is the four-task distribution's odd one out, but it shares the same shape: a deformable object constrained to a small state space.
The novelty is spent fast inside those four task classes. Each new robot running pick-and-place on sheet metal adds another sample from a distribution the field has already largely covered. The Tesla Austin data is the visible failure case for the same argument at much larger scale. IIoT World's aggregation of Figure AI and BMW announcements reports the Spartanburg deployment hit above 99% placement accuracy, 84-second cycle times, and over 90,000 sheet metal parts loaded across 1,250 operational hours over 11 months, a clean example of the four-task pattern rather than a counterexample. The same aggregation reports the Figure 02 fleet contributed to production of 30,000 plus BMW X3 vehicles, the first production-validated deployment metrics for humanoid automation in automotive manufacturing. Xiaomi's reported 98% success rate on a factory humanoid task fits the same mold: a saturable pick-and-place job the field already knows how to do.
High success rates demand constrained problems, and constrained problems don't transfer across a high-dimensional task space. A model that nails sheet-metal pick-and-place is not, in any obvious way, closer to a model that can tidy a kitchen. The novelty pump has to run on each new task, and each task has to be valuable enough to justify the pump.
That is where the borrowing breaks. Robotics startup funding in 2025 passed $8.5 billion, the highest since 2021, with humanoid-specific funding near $4.3 billion, a roughly 6x increase from 2018, per IIoT World's aggregation of Crunchbase and Bank of America figures. The capital is following the self-driving playbook, but the underlying task space is not shaped like a self-driving task space. The four saturable task classes can each justify a focused deployment. They do not justify the claim that the industry is building a general-purpose data flywheel.
The counterpoint is real. Self-driving is the cleanest case where in-domain data really does compound, with a single task, enormous value per mile, and a safety case that forces tight, ongoing collection. Tesla's million-mile-plus data and the still-present Austin safety driver are not a refutation. They are a warning. A flywheel that needs this much mileage, this many years, and this much per-mile value to keep spinning is not the same machine as a humanoid bolting a sheet-metal part in Spartanburg. Treating the two as variations of the same bet is the structural mistake the field is currently making.
The tell for a reader sorting a humanoid announcement: if the demo is one of the four task classes, if the metric is accuracy or cycle time, and if the company says "flywheel," discount the implied generality and look at the unit economics of the task on its own.