DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
DeCoPrune keeps the long-range video memory that matters while cutting more than 85% of historical KV tokens.
The paper frames cache pruning as a denoising-consistency test, keeping tokens whose intermediate clean prediction differs most from the final denoised value. Its premise is that high-discrepancy tokens carry visual evidence the retained context cannot easily predict. The authors introduce CMBench, with 58 roughly one-minute context episodes and 116 continuation tasks built around object or scene recall. On LingBot World v2, they report near-FullKV recall and more than 4x faster continuation generation. HF Daily Papers' note
The paper frames cache pruning as a denoising-consistency test, keeping tokens whose intermediate clean prediction differs most from the final denoised value. Its premise is that high-discrepancy tokens carry visual evidence the retained context cannot easily predict. The authors introduce CMBench, with 58 roughly one-minute context episodes and 116 continuation tasks built around object or scene recall. On LingBot World v2, they report near-FullKV recall and more than 4x faster continuation generation. HF Daily Papers' note
score 4