DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD replaces DMD’s auxiliary student-score model with discriminator heads that learn log-density ratios directly.
The paper frames distribution matching as a classification problem, using two heads on a shared backbone to separate real data and teacher samples from student outputs. At the discriminator optimum, the authors say the resulting losses recover the DMD distribution-matching gradient. They also add gap-based reweighting to adjust teacher supervision by noise level. Reported results include one-step ImageNet-64x64 FID of 1.04 and four-step SDXL COCO-10K FID of 14.47.
HF Daily Papers' note
The paper frames distribution matching as a classification problem, using two heads on a shared backbone to separate real data and teacher samples from student outputs. At the discriminator optimum, the authors say the resulting losses recover the DMD distribution-matching gradient. They also add gap-based reweighting to adjust teacher supervision by noise level. Reported results include one-step ImageNet-64x64 FID of 1.04 and four-step SDXL COCO-10K FID of 14.47.
HF Daily Papers' note
score 5