Researchers combine a training method that runs all timesteps in parallel with a stabilization technique for chaotic dynamics (generalized teacher forcing, or GTF), reporting up to 870x speedups on simulated long chaotic time series and beating
The bottleneck in teaching a neural network to model a chaotic system (weather, turbulence, a heart rhythm, a financial market) was never the math. It was the order in which the data had to be fed in. Recurrent neural networks, the classic sequence-learning model, train step by step, propagating the error signal backward through every step. For a long chaotic time series, that is slow and prone to divergence. A new method by Felix Hess, Sebastian Götz and Tim Durstewitz runs the network across the full sequence at once and stabilizes the optimization with a separate stability trick.
The approach pairs two ingredients. The first is DEER, a parallel-in-time algorithm that computes Newton updates across the whole sequence simultaneously, the way a sparse linear solver handles an entire system of equations rather than walking it row by row. The second is generalized teacher forcing (GTF), a stabilization method the same group introduced at ICML 2023 to keep training gradients bounded when the underlying dynamics are chaotic. Combined, the Hess et al. paper reports that DEER can be run in parallel without the divergence that has historically made parallel-in-time training of recurrent nets unstable on chaotic data.
On simulated benchmarks, the authors report average-case sequence-length scaling of O((log T)^2), so the training cost grows only with the logarithm of the sequence length squared rather than linearly. On the same benchmarks they report speedups up to 870 times over sequential training while matching or improving reconstruction quality. The preprint also reports improved reconstruction on sequences longer than 10,000 steps when the data carries long time scales, and was posted to arXiv on May 12, 2026 as a 29-page preprint.
The benchmark these claims are measured against is dynamical systems reconstruction, or DSR: recovering the rules of a changing system from observed time series. The authors report their method beats current state-of-the-art sequence models, including Mamba and other state space models, on this specific benchmark. The comparison is on the authors' own evaluation, not on independent reproduction, and the underlying systems are simulated. Independent replication, deployment adoption, and measured economic impact are not established by the source.
Two conditions bound the result. The "long sequences" finding applies when the data has long time scales, not as a general claim that parallel-in-time training makes any chaotic time series easier. The paper's critique of linear-time recurrences does not generalize to every state space model or every Mamba application. The 870x upper bound is a benchmark result on a specific task, not a guarantee about training cost across architectures.
The wall was the sequential order of training: the network had to ingest the chaotic series step by step, with the error signal winding back through every step. DEER changes that. The network sees the whole sequence at once, and GTF keeps the parallel optimization from blowing up on the chaotic dynamics. For researchers working on long chaotic time series, on climate, fluid flow, neuroscience, or finance, that combination is a real shift in what can be attempted on a small GPU cluster rather than a supercomputer.
The result lands while sequence-learning research is actively debating how to scale beyond quadratic attention and how to model long, real-world chaotic data. A paper that turns one of those bottlenecks into a sub-quadratic training step is worth watching. The independent question is whether the speedup survives contact with real, noisy data outside the paper's simulated systems, and whether outside groups reproduce the Mamba comparison.