Explicit Layer Modeling for Video Object Insertion and Layer Decomposition
TriLayer gives video-editing models explicit foreground, background, and composite layers to learn from.
The paper introduces a large triplet video dataset with aligned composite, background, and foreground videos, including visual effects tied to the foreground object. Its DBL-Diffusion framework jointly models RGB composites and RGBA foreground layers through shared denoising and cross-branch interaction. The authors apply it to object insertion and layer decomposition, reporting better insertion fidelity and decomposition quality from explicit layer supervision. HF Daily Papers' note
The paper introduces a large triplet video dataset with aligned composite, background, and foreground videos, including visual effects tied to the foreground object. Its DBL-Diffusion framework jointly models RGB composites and RGBA foreground layers through shared denoising and cross-branch interaction. The authors apply it to object insertion and layer decomposition, reporting better insertion fidelity and decomposition quality from explicit layer supervision. HF Daily Papers' note
score 5