Megadose AI progress, ranked and analyzed.

When Does Muon Help Agentic Reinforcement Learning?

· HF Daily Papers ·
Muon’s advantage shows up when RL post-training can use a larger stable effective step.

The paper tests Muon against AdamW on ALFWorld with Qwen2.5 models from 0.5B to 3B. Under the same KL and clipping setup, fan-in Muon stays stable at a more aggressive step size and improves late success at `3e-5` versus an AdamW `1e-6` baseline. The gain is not universal: tuned AdamW nearly catches it in one 3B GraphGPO setting, and matching update RMS removes Muon’s late-success edge. The authors frame the result as an operating-regime finding, not a broad optimizer ranking. HF Daily Papers' note

score 5

Categories: Research