A new preprint argues video models break physics because they see every frame at once. Their fix: rough sketch first, then a parallel cleanup pass — with a LoRA adapter builders can drop into existing video models.
A generated clip shows a hand reaching for a Rubik's Cube. The fingers pass through the plastic and the cube teleports halfway across the table. A different model produces a ball that rolls off a ledge and then hovers in mid-air. These are not render bugs. They are the everyday failures of today's text-to-video systems, and a new preprint from researchers at the University of Cambridge and Toyota Motor Europe argues they are not solved by throwing more data at the problem.
The team, led by Jeffrey Hu with Daniel Olmeda Reino of Toyota Motor Europe and Ayush Tewari of Cambridge, frames the failure as a structural property of bidirectional video diffusion transformers such as Wan. Their argument leans on the Serial Scaling Hypothesis from Liu et al., 2026, which places standard diffusion inside the complexity class TC0, a regime with bounded sequential depth. In plain terms, a denoiser that sees every frame at once can paint each region with locally plausible detail but cannot reason across time to keep those details consistent. Gravity reverses. Identities drift. Hands pass through cups.
The proposed remedy, called S2PD for Serial-to-Parallel Diffusion, splits the sampling schedule in two. At high noise, the model runs autoregressively, generating frames in order so each step conditions on what came before. Once the noise is low enough that the global structure is mostly settled, the model switches to parallel denoising, refining every frame simultaneously. The intuition is borrowed from how a rough sketch is laid down before a final cleanup pass: get the timeline right, then polish the picture.
The team ships two implementations. The first is a pixel-space diffusion transformer trained from scratch. The second is a LoRA adapter with causal attention bolted onto an existing pretrained video model, the more practical route for anyone who already runs Wan-class weights. The hybrid recipe is also faster than running the model in fully serial mode, because only the high-noise phase pays the autoregressive cost.
Evaluation spans four settings that cover different kinds of rules. On the procedurally generated Kubric MOVi-A and MOVi-C benchmarks, which simulate rigid-body physics, S2PD follows gravity and collision rules more reliably than matched bidirectional baselines. On MPMWorlds, which models soft materials, the same pattern holds for plausible deformation. On a real captured dataset of a hand manipulating a Rubik's Cube, the hybrid schedule produces more temporally stable object interactions. The paper also contributes new metrics aimed specifically at counting rule violations and physical errors, filling a gap prior video benchmarks left open.
The training data is procedurally generated and in distribution, so the result is a lower bound: even when the model has effectively unlimited examples of correct physics, the bidirectional architecture still fails. The fix improves over matched baselines and over fully serial sampling, but it does not claim general video generation has been solved. Open-world footage, long horizons, and rare interactions are not addressed. The serial phase is also slower than purely parallel diffusion, so the sampling speedup is relative to other serial methods, not to the wider field.
The release, funded by Toyota Motor Europe and the UK's Isambard-AI National AI Research Resource, includes code, datasets, and model weights. Outside labs can therefore reproduce the claim on their own video models. The contribution is a specific architectural lever, a hybrid schedule plus a LoRA adapter, that lets practitioners test whether causality can be taught to video models without retraining from scratch.