Leju built a domain specific robot model from 600+ hours of its own hardware data, scoring 48.27% on a 25 task benchmark and leading three general AI peers, but most tasks failed.
A Chinese humanoid robotics startup is making the first named bet that the missing layer in factory AI is "mid-training" on its own robots' data, and using a 48.27% success rate on a 25-task industrial benchmark to argue for it. Leju Robotics (乐聚机器人) calls the model KUAVO VLA and positions it as a middle layer between a general foundation model and a humanoid's task-specific skills.
"Mid-training" here means something specific. Modern robot foundation models, called vision-language-action (VLA) models, are pretrained on broad datasets so a humanoid can understand images, language, and motion in general terms. That is the "college" layer. Per-skill training tunes the model for one task, like sorting parts. Mid-training sits between them: the vendor takes a general VLA, then trains it on hundreds of hours of its own robot's body, in its own factories, so the model learns the specific shape, joint limits, and sensor quirks of one machine. The 600+ hours Leju used for KUAVO VLA are all recorded on the company's KUAVO humanoid body.
Leju ran a 25-task benchmark, 20 industrial tasks the company designed plus 5 from the public GM100 set, and put KUAVO VLA against three general-purpose VLA peers: π 0.5 from Physical Intelligence, NVIDIA's GR00T N1.7, and Lingbot-VLA 2.0. KUAVO VLA scored 48.27% task success and 74.51% on a process score, ranking first on both. Versus Lingbot-VLA 2.0, the gap is 32 percentage points on success and 32.45 on process.
"First" in a four-model comparison is real, and it is thin. 48.27% means more than half of the 25 industrial tasks still failed, even with the body-matched training. The leiphone exclusive, which carried the launch, also notes that KUAVO VLA hits 100% success on "some tasks" and roughly doubles efficiency in many scenarios; both claims are qualified in the source. None of the three named peers is independently broken out at the per-task level in the article, and the strongest of the three, π 0.5, has an open paper but was not the one chosen for the head-to-head number. Lingbot-VLA 2.0 is.
That baseline selection matters. Leju's exclusive puts the strongest detailed numbers against the weakest-named peer, which is the standard pattern for company-asserted benchmark wins, and the 32-point delta is real but is one delta against one model in a self-defined test. The other two peers are not numerically broken out, so the "first place" headline collapses several asymmetries into one ranking. The 48.27% absolute number is the more honest read of where industrial embodied AI is right now.
Same-form-factor real-robot data is what lets a small vendor catch up to large foundation models; it is also the textbook overfit risk. The model learns the KUAVO body, and the 25 industrial tasks, and possibly the lab lighting. None of the source material shows transfer to a different humanoid, to a different factory cell, or to a non-industrial domain. The body of evidence is consistent with the claim, not a verification of it. Until a third party runs the same benchmark on a different robot, the "mid-training" thesis is plausible rather than proven.
The launch is paired with platform and ecosystem work, which is where the angle becomes a real industrial bet rather than a model release. Leju is publishing developer tooling on GitHub, running a KUAVO data challenge, and has tied its industrial pipeline to a Huawei Cloud Embodied Intelligence Industry Innovation Center in Shenzhen. The corporate news landing page is the first-party surface for the launch. The framing is consistent: a model plus a developer surface plus a cloud partner, sold as a stack to industrial customers.
Leju EVP Ke Zhendong (柯真东) told leiphone that "future embodied developers will need to dive into specific industries and pull out the right skills in context. For them, Leju is offering the platform, the model, and the toolchain," in a translation of the original Chinese remark. The quote is on the developer-ecosystem lane, not on the benchmark.
The falsifier is whether the body-binding survives a hardware change. If KUAVO VLA's 48.27% on KUAVO hardware is matched, or even cut in half, on a different humanoid, the mid-training thesis is real. If it collapses to near zero off the KUAVO body, Leju has built a model that learned one robot, not industrial AI. Until the company, or a partner, releases numbers on a second form factor, the launch is a credible architectural argument with a thin self-defined benchmark behind it.