Megadose AI progress, ranked and analyzed.

MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining

· ArXiv · AI/CL/LG ·
MeqMuon changes Muon by balancing update matrices across both rows and columns while dropping AdamW-style second-moment storage.

The paper frames row-only normalization as too narrow for the imbalance patterns that appear in LLM pretraining updates. MeqMuon applies matrix equilibration so the normalization can adapt without manual tuning. The authors report better convergence than AdamW, Muon, and other baselines, with lower optimizer-state memory use. ArXiv · AI/CL/LG's note

score 5

Categories: Research