Improved Distributional Diffusion Models
The paper says DDMs can be made practical for ImageNet-scale few-step generation without distillation.
The authors defer particle expansion to later transformer layers to reduce multi-particle training cost. They also add time-dependent scoring-rule schedules instead of using one fixed trade-off across sampling budgets. In a DiT-based latent setup, the model reports 4.48 FID at 4 steps and 2.38 at 50 steps on class-conditional ImageNet-256, trained from scratch in one stage. The abstract says the same recipe transfers to text-to-image generation. HF Daily Papers' note
The authors defer particle expansion to later transformer layers to reduce multi-particle training cost. They also add time-dependent scoring-rule schedules instead of using one fixed trade-off across sampling budgets. In a DiT-based latent setup, the model reports 4.48 FID at 4 steps and 2.38 at 50 steps on class-conditional ImageNet-256, trained from scratch in one stage. The abstract says the same recipe transfers to text-to-image generation. HF Daily Papers' note
score 4