MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining
MeqMuon changes Muon by balancing update matrices across both rows and columns while dropping AdamW-style second-moment storage.
The paper frames row-only normalization as too narrow for the imbalance patterns that appear in LLM pretraining updates. MeqMuon applies matrix equilibration so the normalization can adapt without manual tuning. The authors report better convergence than AdamW, Muon, and other baselines, with lower optimizer-state memory use. ArXiv · AI/CL/LG's note
The paper frames row-only normalization as too narrow for the imbalance patterns that appear in LLM pretraining updates. MeqMuon applies matrix equilibration so the normalization can adapt without manual tuning. The authors report better convergence than AdamW, Muon, and other baselines, with lower optimizer-state memory use. ArXiv · AI/CL/LG's note
score 5