Megadose Built for builders and researchers.

On the Principles Behind Neural Network Optimizers

· ArXiv · AI/CL/LG ·
Adam’s behavior hinges on hyperparameter regimes, and the paper argues its edge on Transformers comes from Hessian structure.

Yushun Zhang’s thesis revisits Adam’s convergence debate, claiming a problem-dependent phase transition where batch-size-aware settings converge but small-β₂ regimes can diverge. It says Transformer training drives the Hessian toward a near-block-diagonal form with strong block heterogeneity, making Adam’s diagonal preconditioner effective. The work traces that structure to repeated multiplication of large matrix variables and uses random matrix theory to analyze it. Those findings motivate Adam-mini, presented as cutting Adam’s memory use by 50% while preserving performance. ArXiv · AI/CL/LG's note

score 5

Categories: Research