ProAct, a vision language action model, anchors each action chunk in the robot's last motion and a predicted future frame. Reported as an arXiv preprint, not yet peer reviewed.
A new robot policy, ProAct, decides each chunk of action by anchoring its guess in the robot's last motion and a predicted future frame, instead of starting from a generic noise prior. The result, reported in an arXiv preprint this week, is roughly 50% fewer denoising steps, up to 25.8% lower inference latency, and up to 34.8% higher throughput against the pi-0.5 baseline.
The work comes from Harbin Institute of Technology (Shenzhen), ZTE, and Sun Yat-sen University, and reframes a Vision-Language-Action model as a flow-based sampler that no longer sprays noise and hopes for the best. Two extra modules, a Proposal Expert that extrapolates recent motion and a World Expert built on a frozen V-JEPA 2 target encoder, give the sampler a motion-respecting starting point. The project page and full paper report 98.4% on LIBERO, 86.8% on LIBERO-Plus, 60.3% on a randomized hard subset of RoboTwin 2.0, and real-robot task success of 78.7% on a GALAXEA R1 Lite and 74.7% on an AgileX Cobot Magic.
For builders, fewer internal steps per motion chunk means higher control rates and lower latency on the same hardware, the kind of plumbing that turns demo reels into deployed systems.
What is still unknown: this is a v1 preprint with no peer review, the headline efficiency numbers are measured against a single baseline family, and reproducibility of the V-JEPA 2 stack is untested here.