Megadose AI progress, ranked daily.

Effective Learning Rate Governs Loss Dynamics in Language Model Pretraining

· ArXiv · AI/CL/LG ·
Matching the ratio of learning rate to parameter norm made very different pretraining runs follow nearly the same loss curve.

The paper calls that ratio the effective learning rate, or ELR. Its experiments report collapse errors typically on the order of a few thousandths across optimizers, architectures, datasets, and scales. The authors also find that weight decay and Hyperball affect loss mainly through the ELR schedules they create. They argue ELR gives a common coordinate for learning-rate schedules, norm control, and delayed acceleration in pretraining. ArXiv · AI/CL/LG's note

score 6

Categories: Research