Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
SAMuon cuts Muon’s token budget by 13.3% to 24.0% at the same validation loss.
The paper argues Muon’s gain comes from how it allocates steps across singular directions in Transformer training. Its probes find a small, volatile spectral “head” that needs conservative steps and a larger “bulk” that can take bigger ones. The authors say Muon’s uniform scaling leaves that bulk underused, then propose SAMuon and SAMuon-lite to amplify it without persistent extra optimizer state. In modded-nanogpt runs from 124M to 1B parameters, both beat tuned AdamW and Muon baselines across the tested scale and batch settings. ArXiv · AI/CL/LG's note
The paper argues Muon’s gain comes from how it allocates steps across singular directions in Transformer training. Its probes find a small, volatile spectral “head” that needs conservative steps and a larger “bulk” that can take bigger ones. The authors say Muon’s uniform scaling leaves that bulk underused, then propose SAMuon and SAMuon-lite to amplify it without persistent extra optimizer state. In modded-nanogpt runs from 124M to 1B parameters, both beat tuned AdamW and Muon baselines across the tested scale and batch settings. ArXiv · AI/CL/LG's note
score 6