Muon Meets Mamba: Spectral Optimization for State Space Models
Muon’s gain on Mamba-2 shows up when it trains the output projection, not broadly across projections.
The paper compares Muon with AdamW on Mamba-2 130M while changing only which weight groups use Muon. Its reported advantage is token efficiency, holding across two corpora and two token budgets and remaining after training beyond the compute-optimal point. Applying Muon to the output projection alone beats using it on the input projection or on both. The authors say conditioning is not the explanation: Muon lowers condition numbers where applied, but the better-conditioned input projection is not where the benefit appears. ArXiv · AI/CL/LG's note
The paper compares Muon with AdamW on Mamba-2 130M while changing only which weight groups use Muon. Its reported advantage is token efficiency, holding across two corpora and two token budgets and remaining after training beyond the compute-optimal point. Applying Muon to the output projection alone beats using it on the input projection or on both. The authors say conditioning is not the explanation: Muon lowers condition numbers where applied, but the better-conditioned input projection is not where the benefit appears. ArXiv · AI/CL/LG's note
score 5