From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes
Learning-rate and batch-size schedules can change whether pre-training loss keeps a clean power law.
The paper studies noisy online SGD with linear random features and separates loss dynamics into a forcing term and a memory term. It argues that power laws come from cumulative low-spectrum mass, not necessarily from each eigenvalue or target coefficient following its own power law. Under joint schedules, intrinsic time tracks optimization progress, while `B/η` governs noise injection, producing regimes where schedules preserve, alter, or break scaling behavior. Controlled nanoGPT experiments support the schedule equivalence and surrogate-loss claims.
ArXiv · AI/CL/LG's note
The paper studies noisy online SGD with linear random features and separates loss dynamics into a forcing term and a memory term. It argues that power laws come from cumulative low-spectrum mass, not necessarily from each eigenvalue or target coefficient following its own power law. Under joint schedules, intrinsic time tracks optimization progress, while `B/η` governs noise injection, producing regimes where schedules preserve, alter, or break scaling behavior. Controlled nanoGPT experiments support the schedule equivalence and surrogate-loss claims.
ArXiv · AI/CL/LG's note
score 5