One Editor, Many Edits: A Unified Training-Free Framework for Diverse Video Editing
EditVid claims one training-free video editor can handle both instruction-led and reference-led edits.
The framework combines sparse causal memory, post-attention token injection, and soft latent blending to preserve coherence, identity, and edit locality. It covers style transfer, attribute changes, object insertion, part-level edits, and subject replacement. On FiVE, it reports 78.16 FiVE-Acc versus 58.95 for the strongest evaluated training-free baseline, with competitive IVEBench results. A user study preferred EditVid overall 51.8% of the time against seven competing methods. ArXiv · AI/CL/LG's note
The framework combines sparse causal memory, post-attention token injection, and soft latent blending to preserve coherence, identity, and edit locality. It covers style transfer, attribute changes, object insertion, part-level edits, and subject replacement. On FiVE, it reports 78.16 FiVE-Acc versus 58.95 for the strongest evaluated training-free baseline, with competitive IVEBench results. A user study preferred EditVid overall 51.8% of the time against seven competing methods. ArXiv · AI/CL/LG's note
score 4