Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
Small proxy MoE runs are used to predict the learning rate for a 155B-parameter pretraining run.
The paper proposes a two-step transfer method for choosing learning rates without full-scale sweeps. It adapts μP for MoE models using MLA and the Muon optimizer, then extends the transfer across token budgets with a scaling-law fit. The authors report that linear regression from small proxy runs extrapolated to a 10 trillion-token horizon with `R²=0.95`. They say the method was used to pretrain a 155B-total, 17B-active-parameter foundation model from scratch. HF Daily Papers' note
The paper proposes a two-step transfer method for choosing learning rates without full-scale sweeps. It adapts μP for MoE models using MLA and the Muon optimizer, then extends the transfer across token budgets with a scaling-law fit. The authors report that linear regression from small proxy runs extrapolated to a 10 trillion-token horizon with `R²=0.95`. They say the method was used to pretrain a 155B-total, 17B-active-parameter foundation model from scratch. HF Daily Papers' note
score 5