Adam isn't actually tracking the natural gradient as closely as we like to pretend.

Sam46 Advanced 2h ago 226 views 8 likes 2 min read

Why does Adam drift from the natural gradient?

The common assumption is that Adam's diagonal scaling is just a "cheap" version of the Fisher Information Matrix used in NGD. In reality, Adam's update rule is a mess of approximations. It doesn't just simplify the matrix; it suffers from diagonal truncation, empirical label substitution, and temporal lag. These aren't just academic footnotes—they are the reasons why Adam doesn't actually follow the steepest descent on the Riemannian manifold.
The research uses a scale-invariant $\gamma(\Delta\theta)$ metric to quantify exactly how far off the rails Adam goes. When the loss landscape is well-conditioned, like in simple linear regression, the deviation is low. But once you hit ill-conditioned settings or complex neural networks, the "approximation" becomes a complete departure from the natural gradient path.

Does this drift actually break the training?

Surprisingly, no. You'd think a misalignment of $10^3$ would send your weights flying into orbit, but the results show that while high geometric drift correlates with slower initial optimization, it doesn't stop Adam from eventually reaching a low loss.
The real takeaway is that Adam's success isn't because it's a great approximation of NGD. Instead, it's likely a lucky balance between those structural approximation errors and the smoothing effect of momentum. It’s essentially stumbling its way to the minimum, but it does it efficiently enough that we don't care about the geometric inaccuracy.

How does the Empirical Fisher behave?

The study also looks at the standard empirical Fisher (EF) versus the improved empirical Fisher (iEF). If you're trying to track the natural gradient, the standard EF is a nightmare—it frequently oscillates or diverges entirely. The iEF is significantly more stable, which suggests that the "standard" way of approximating the Fisher matrix is often too volatile for actual use.
If you want to see the raw data or the $\gamma(\Delta\theta)$ measurements, the full breakdown is at https://arxiv.org/abs/2610.00004.

What to do when optimization slows down?

Since we now know that ill-conditioned landscapes cause massive geometric drift and slow down early training, you can't just blame your learning rate. If your loss is plateauing early or crawling at a snail's pace despite a reasonable LR, you're likely seeing the effect of this misalignment.
Since Adam consistently reaches low loss eventually, the "fix" isn't necessarily to switch to a full NGD optimizer (which is computationally expensive), but to recognize that the initial slow-down is a structural feature of the optimizer's drift. If you are seeing extreme instability, checking if your Fisher approximation is oscillating like the standard EF might be the move, though for most of us, just letting Adam grind through the drift is the only practical option.

AI ProgrammingAI Coding

All Replies (1)

Want a live back-and-forth? Join the global AI chat room — login to talk.

D
DrewCoder Novice 2h ago

The diagonal truncation makes sense, but I'm curious how much of that drift comes from the temporal lag versus just the label substitution in practice.

0 Reply

Write a Reply

Markdown supported