CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
CoinVE-200K is built for video edits that combine several instructions in the same clip.
The dataset includes 1080p video-editing pairs up to 201 frames, with each sample combining 2 to 5 atomic edits. Those edits cover humans, objects, and backgrounds, including addition, removal, modification, and stylization. The paper also introduces CoinVE-Bench and a 22B model, CoinVE-Edit, which the authors say performs strongly on instruction following, compositional accuracy, visual quality, and temporal consistency. HF Daily Papers' note
The dataset includes 1080p video-editing pairs up to 201 frames, with each sample combining 2 to 5 atomic edits. Those edits cover humans, objects, and backgrounds, including addition, removal, modification, and stylization. The paper also introduces CoinVE-Bench and a 22B model, CoinVE-Edit, which the authors say performs strongly on instruction following, compositional accuracy, visual quality, and temporal consistency. HF Daily Papers' note
score 4