Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
Mask Forcing targets the mode collapse behind washed-out autoregressive video diffusion outputs.
The paper says DMD-trained AR video students can drift into over-saturated, over-smoothed generations because reverse KL pushes them toward too few teacher modes. Its fix is a dual-noise masking rollout: during student self-rollout, random spatial and temporal masks inject cleaner signals into noisier inputs. The authors argue those cleaner tokens guide denoising, reduce error accumulation, and make the student explore more of the teacher distribution. They report improvements across multiple AR video diffusion distillation methods without real video data or extra post-training. HF Daily Papers' note
The paper says DMD-trained AR video students can drift into over-saturated, over-smoothed generations because reverse KL pushes them toward too few teacher modes. Its fix is a dual-noise masking rollout: during student self-rollout, random spatial and temporal masks inject cleaner signals into noisier inputs. The authors argue those cleaner tokens guide denoising, reduce error accumulation, and make the student explore more of the teacher distribution. They report improvements across multiple AR video diffusion distillation methods without real video data or extra post-training. HF Daily Papers' note
score 4