FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
FlowMimic claims it can train video editing from image-editing samples generated into paired video data in real time.
The paper proposes a pixel-pair temporal warped flow field to create corresponding video editing samples without object masks or multi-stage curation. It treats images as a special case of video and adds mimic losses meant to align generation and editing behavior across the two modalities. The authors also add sense-related tasks and region-aware losses so the model learns where an instruction applies without needing an explicit mask at inference. HF Daily Papers' note
The paper proposes a pixel-pair temporal warped flow field to create corresponding video editing samples without object masks or multi-stage curation. It treats images as a special case of video and adds mimic losses meant to align generation and editing behavior across the two modalities. The authors also add sense-related tasks and region-aware losses so the model learns where an instruction applies without needing an explicit mask at inference. HF Daily Papers' note
score 5