Megadose AI progress, ranked and analyzed.

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

· ArXiv · AI/CL/LG ·
The paper claims a single scalar separates stable and unstable effective learning-rate regimes under normalization and weight decay.

Amin, Chang, and Khanna argue that scale invariance creates a feedback loop between learning-rate schedules, weight decay, and parameter norms. Their discrete-time law predicts when effective steps contract or expand, with norm growth acting as a self-quenching force. In a solved normalized regression model, the balance point is unstable, so constant learning rate plus weight decay produces recurrent dynamics rather than a stable interior equilibrium. Tests across small models and GPT-2 settings reportedly match the predicted boundary closely, with performance peaking near it. ArXiv · AI/CL/LG's note

score 5

Categories: Research