RECAP-Forcing: Retaining Content Appearances for Long Video Generation
The paper’s key move is to keep memory slots for newly visible appearances, not just recent frames.
RECAP-Forcing treats long video consistency as a memory-selection problem under a finite attention window. It retains KV cache entries when subjects, objects, disoccluded regions, or scenes first appear, so identities can be recalled later. The method combines an initial attention sink with an optical-flow-based novelty bank. The authors say it is training-free, adds no learnable parameters, and improves quality and semantic fidelity across multiple baselines. HF Daily Papers' note
RECAP-Forcing treats long video consistency as a memory-selection problem under a finite attention window. It retains KV cache entries when subjects, objects, disoccluded regions, or scenes first appear, so identities can be recalled later. The method combines an initial attention sink with an optical-flow-based novelty bank. The authors say it is training-free, adds no learnable parameters, and improves quality and semantic fidelity across multiple baselines. HF Daily Papers' note
score 5