Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
SparsePR keeps video-model attention quality while running only about a quarter of the attention pairs.
The paper proposes a training-free sparse attention method for video generation and world models. It pairs response-coupled routing with a small exact-query probe that fits an affine correction for the residual left by skipped attention interactions. Across four video and world models, the authors report lower attention-reconstruction error and preserved generation quality at 22.0-26.0% executed-pair density. Reported end-to-end speedups range from 1.48x to 2.61x. HF Daily Papers' note
The paper proposes a training-free sparse attention method for video generation and world models. It pairs response-coupled routing with a small exact-query probe that fits an affine correction for the residual left by skipped attention interactions. Across four video and world models, the authors report lower attention-reconstruction error and preserved generation quality at 22.0-26.0% executed-pair density. Reported end-to-end speedups range from 1.48x to 2.61x. HF Daily Papers' note
score 5