Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
SplitMoE replaces uniform expert balancing with separate semantic and generic experts for video diffusion.
The paper argues that standard token-wise MoE routing fragments coherent video patches because it pushes expert use toward uniformity. SplitMoE instead lets semantic experts group tokens by higher-level attributes while generic experts preserve residual visual detail and generation capacity. The authors report faster convergence, more coherent routing, and better video generation quality than load-balanced MoEs under the same activated-parameter budget. It was accepted as a NeurIPS 2026 Spotlight paper. HF Daily Papers' note
The paper argues that standard token-wise MoE routing fragments coherent video patches because it pushes expert use toward uniformity. SplitMoE instead lets semantic experts group tokens by higher-level attributes while generic experts preserve residual visual detail and generation capacity. The authors report faster convergence, more coherent routing, and better video generation quality than load-balanced MoEs under the same activated-parameter budget. It was accepted as a NeurIPS 2026 Spotlight paper. HF Daily Papers' note
score 5