A Solvable Model of Adaptive Learning Rate Rescaling: Acceleration, Stability & Scaling
Normalized SGD can speed learning early, but the same rescaling later pushes training toward stability limits.
The paper studies fixed-norm updates in a random-feature model, where the effective learning rate rises as gradients shrink. Its theory predicts an initial acceleration, then a finite-step-size breakdown into marginal stability. The late-time regimes determine when wider models or larger batches actually reduce serial training time at comparable compute. Linearized ResNet experiments on CIFAR-5M are reported as support for those acceleration, breakdown, and scaling predictions. ArXiv · AI/CL/LG's note
The paper studies fixed-norm updates in a random-feature model, where the effective learning rate rises as gradients shrink. Its theory predicts an initial acceleration, then a finite-step-size breakdown into marginal stability. The late-time regimes determine when wider models or larger batches actually reduce serial training time at comparable compute. Linearized ResNet experiments on CIFAR-5M are reported as support for those acceleration, breakdown, and scaling predictions. ArXiv · AI/CL/LG's note
score 5