MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
MC-Sparse targets the specific failure modes that make sparse attention lose quality at high sparsity.
The paper traces the gap to token-grouping constraints, bad interaction selection, and discarded attention contributions. Its method selects individual KV tokens, groups similar queries for GPU efficiency, and caches metadata plus dense-sparse residuals across later denoising steps. The authors report better fidelity than prior sparse-attention baselines, with no visible quality degradation in their video and 3D generation tests. They cite 1.80x denoising speedup on Minimax-H3-Base and 2.32x on 3D asset generation versus dense attention. HF Daily Papers' note
The paper traces the gap to token-grouping constraints, bad interaction selection, and discarded attention contributions. Its method selects individual KV tokens, groups similar queries for GPU efficiency, and caches metadata plus dense-sparse residuals across later denoising steps. The authors report better fidelity than prior sparse-attention baselines, with no visible quality degradation in their video and 3D generation tests. They cite 1.80x denoising speedup on Minimax-H3-Base and 2.32x on 3D asset generation versus dense attention. HF Daily Papers' note
score 4