Megadose AI progress, ranked and analyzed.

Muon Meets Mamba: Spectral Optimization for State Space Models

· ArXiv · AI/CL/LG ·
Muon’s gain on Mamba-2 shows up when it trains the output projection, not broadly across projections.

The paper compares Muon with AdamW on Mamba-2 130M while changing only which weight groups use Muon. Its reported advantage is token efficiency, holding across two corpora and two token budgets and remaining after training beyond the compute-optimal point. Applying Muon to the output projection alone beats using it on the input projection or on both. The authors say conditioning is not the explanation: Muon lowers condition numbers where applied, but the better-conditioned input projection is not where the benefit appears. ArXiv · AI/CL/LG's note

score 5

Categories: Research