Megadose AI progress, ranked and analyzed.

From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

· ArXiv · AI/CL/LG ·
Learning-rate and batch-size schedules can change whether pre-training loss keeps a clean power law.

The paper studies noisy online SGD with linear random features and separates loss dynamics into a forcing term and a memory term. It argues that power laws come from cumulative low-spectrum mass, not necessarily from each eigenvalue or target coefficient following its own power law. Under joint schedules, intrinsic time tracks optimization progress, while `B/η` governs noise injection, producing regimes where schedules preserve, alter, or break scaling behavior. Controlled nanoGPT experiments support the schedule equivalence and surrogate-loss claims.

ArXiv · AI/CL/LG's note

score 5

Categories: Research