CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization
CMuon targets a DiT optimizer failure caused by fused weight tensors.
The paper says standard Muon struggles late in DiT training because fused AdaLN and QKV-style tensors cause implicit subspace coupling during orthogonalization. CMuon splits those matrices into independent chunks before applying the optimizer step. In experiments, a 675M-parameter DiT reached 1.18 FID on ImageNet 256 after 200 epochs, which the authors report as more than a 2x speedup over AdamW. ArXiv · AI/CL/LG's note
The paper says standard Muon struggles late in DiT training because fused AdaLN and QKV-style tensors cause implicit subspace coupling during orthogonalization. CMuon splits those matrices into independent chunks before applying the optimizer step. In experiments, a 675M-parameter DiT reached 1.18 FID on ImageNet 256 after 200 epochs, which the authors report as more than a 2x speedup over AdamW. ArXiv · AI/CL/LG's note
score 5