DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation
DistillAlign argues that video distillation fails when the student learns sharp outputs but loses the teacher’s distributional coverage.
The paper says common DMD-based autoregressive video distillation splits initialization and refinement in ways that target different distributions. Its evaluation protocol compares student and teacher precision and coverage in a shared latent space, showing gaps that visual scores can miss. The proposed joint distillation pairs DMD’s mode-seeking loss with a consistency distillation constraint meant to preserve coverage and diversity. In experiments, it reports better quality, coverage, and diversity, including outperforming baselines refined with Wan-14B while using a Wan-1.3B DMD teacher. HF Daily Papers' note
The paper says common DMD-based autoregressive video distillation splits initialization and refinement in ways that target different distributions. Its evaluation protocol compares student and teacher precision and coverage in a shared latent space, showing gaps that visual scores can miss. The proposed joint distillation pairs DMD’s mode-seeking loss with a consistency distillation constraint meant to preserve coverage and diversity. In experiments, it reports better quality, coverage, and diversity, including outperforming baselines refined with Wan-14B while using a Wan-1.3B DMD teacher. HF Daily Papers' note
score 4