Megadose AI progress, ranked and analyzed.

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

· ArXiv · AI/CL/LG ·
MALT adds two-sided diagonal preconditioning to Muon to make its updates less sensitive to curvature anisotropy.

The paper says Muon already tackles gradient anisotropy through Newton-Schulz orthogonalization, but can still miss curvature geometry in the loss landscape. MALT preconditions the momentum, orthogonalizes it, then maps it back, with norm grafting used to control update size. The authors also introduce MALTER, which adds adaptive stepsize rescaling for stochastic gradient noise. In GPT-2 Small, Medium, and Large pretraining experiments, they report better results than Muon with nearly the same memory footprint and wall-clock time. ArXiv · AI/CL/LG's note

score 5

Categories: Research