When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay
The paper claims a single scalar separates stable and unstable effective learning-rate regimes under normalization and weight decay.
Amin, Chang, and Khanna argue that scale invariance creates a feedback loop between learning-rate schedules, weight decay, and parameter norms. Their discrete-time law predicts when effective steps contract or expand, with norm growth acting as a self-quenching force. In a solved normalized regression model, the balance point is unstable, so constant learning rate plus weight decay produces recurrent dynamics rather than a stable interior equilibrium. Tests across small models and GPT-2 settings reportedly match the predicted boundary closely, with performance peaking near it. ArXiv · AI/CL/LG's note
Amin, Chang, and Khanna argue that scale invariance creates a feedback loop between learning-rate schedules, weight decay, and parameter norms. Their discrete-time law predicts when effective steps contract or expand, with norm growth acting as a self-quenching force. In a solved normalized regression model, the balance point is unstable, so constant learning rate plus weight decay produces recurrent dynamics rather than a stable interior equilibrium. Tests across small models and GPT-2 settings reportedly match the predicted boundary closely, with performance peaking near it. ArXiv · AI/CL/LG's note
score 5