A new measurement study shows why one Adam run (the default neural network training optimizer), in a neural network under ill conditioning (a sharply uneven loss surface), can drift a thousandfold from the theoretical ideal and still reach the
The Adam run that strays farthest from the theoretical ideal path also reaches the lowest final loss. A new preprint explains why that result, which looks like a paradox, is in fact a feature of how the workhorse engine of modern deep learning actually works.
Adam is the default algorithm that updates a neural network's weights during training. Almost every large model built in the last seven years, from text generators to image classifiers, was trained by some version of it. The theoretical ideal it is usually compared to is natural gradient descent, a more expensive procedure that adjusts weights using the local geometry of the loss landscape rather than a fixed learning rate. The new paper asks a simple question: how close does Adam actually get to that ideal, and when does it miss?
The authors, posting as arXiv:2610.00004, treat Adam's full update rule, including its momentum term, as a diagonal approximation to the empirical Fisher information matrix. That framing decomposes Adam's errors into three sources. The first is diagonal truncation, the loss of the off-diagonal terms that would otherwise couple the weights. The second is empirical label substitution, replacing the true labels of a clean loss with the model's current outputs to build a practical estimate. The third is temporal lag, the delay between the curvature estimate and the current position of the optimizer on the loss surface. None of these is novel on its own. What is new is the yardstick the authors use to compare the resulting update direction against the true natural gradient direction.
The metric is a scale-invariant geometric deviation the paper calls γ(δθ). In plain terms, it measures how far the direction Adam actually takes points away from the direction the natural gradient would have taken, normalized so the result does not blow up as the loss surface changes shape. The authors then run Adam on four controlled loss landscapes: a well-conditioned linear regression, an ill-conditioned linear regression (a setting where the terrain is sharply uneven in one direction and flat in another), a logistic classification problem, and a non-convex multi-layer neural network.
The empirical results are sharp and partly counterintuitive. Under well-conditioning, Adam's geometric deviation from the natural gradient stays low. Under ill-conditioning, the deviation climbs to roughly 10³ by step 300 in the neural network experiment, the highest tracking error the authors observe. That is a thousandfold directional misalignment between what Adam did and what the theoretically ideal path would have done. Yet the same run that hits that thousandfold deviation also reaches the lowest final loss of any setting tested. Geometric drift, in other words, correlates with slower initial optimization but does not degrade the final objective.
The paper's stated implication is that Adam's practical power may come from a balance of structural approximation errors plus the smoothing effect of momentum, not from close tracking of the natural gradient path. The drift is not a bug. It is part of the mechanism that lets the algorithm find good solutions on rough terrain.
There is a falsifier in the results, and it is worth naming. The authors compare two ways of building the practical Fisher estimate: the standard empirical Fisher (EF), which uses the model's own outputs, and an improved version (iEF) that the paper constructs to reduce the bias. The improved variant traces more stable paths. The standard one frequently oscillates or diverges. That contrast is where future optimizer design has room to move, because it isolates a specific component, the empirical label substitution step, that breaks the natural-gradient analogy in a measurable way.
The work is a preprint, not peer-reviewed, and the experimental scope is small: four loss landscapes, controlled rather than at the scale of a frontier model run. The roughly 10³ misalignment figure is specific to the neural network setting and should not be generalized to all ill-conditioned problems without further replication.
What the paper does contribute is a usable measurement. Before γ(δθ), the field had no scale-invariant way to ask how far Adam is from the natural gradient. It does now. That is how better training engines get designed: not by declaring the workhorse broken, but by handing the next round of optimizer researchers a sharper map of the terrain.