A single preprint trains one network to step forward and backward; the gap between where it lands and where it started flags drift in long AI video and physics simulations.
A single AI model can grade its own long rollouts by stepping both forward and backward in time. No second model, no held-out data, no governing equations. The size of the round-trip gap is the error signal.
That is the claim in a new single-author preprint, Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors (arXiv:2608.00675). The same network is trained with a direction flag, so it steps a dynamical system forward or backward in time instead of needing two specialists. Roll forward i steps, then backward i steps; the system should land at its start, and the gap (C_i) is a self-supervised proxy for rollout error.
The authors test the trick on three domains: compressible magnetohydrodynamics, an astrophysical turbulent radiative mixing layer, and CelebV-HQ face videos. On held-out MHD trajectories, C_i ranks rollout error with Spearman correlation 0.91-0.98 at fixed depth and 0.69 ± 0.16 within a single trajectory, the abstract reports, roughly a third weaker within a run than across depths.
The paper lands as long AI video and physics digital twins stretch past the point where deployment-time ground truth exists. The limits: this is a single preprint with author-released code, no independent reproduction, and the "beats two specialist models" claim rests on the author's own benchmarks.