One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
EditVid claims one training-free system can handle both instruction- and reference-guided video edits without separate models.
The paper combines sparse causal memory, post-attention token injection, and soft latent blending to keep edits coherent, preserve identity, and limit changes to the intended areas. It covers style transfer, attribute changes, object insertion, part-level editing, and subject replacement. On FiVE, it reports 78.16 FiVE-Acc versus 58.95 for the strongest evaluated training-free baseline, with competitive IVEBench results. A user study found a 51.8% overall preference for EditVid over seven competing methods. HF Daily Papers' note
The paper combines sparse causal memory, post-attention token injection, and soft latent blending to keep edits coherent, preserve identity, and limit changes to the intended areas. It covers style transfer, attribute changes, object insertion, part-level editing, and subject replacement. On FiVE, it reports 78.16 FiVE-Acc versus 58.95 for the strongest evaluated training-free baseline, with competitive IVEBench results. A user study found a 51.8% overall preference for EditVid over seven competing methods. HF Daily Papers' note
score 5