Megadose AI progress, ranked daily.

On-Policy Self-Distillation in Diffusion Models

· HF Daily Papers ·
DiffusionOPSD turns image-level reward signals into intermediate denoising targets, then refreshes them on-policy during training.

The paper says endpoint rewards are too coarse to tell a diffusion model how each denoising step should change. Its method uses a frozen behavior policy to generate trajectories, builds bounded positive and negative targets from reward gradients, and trains a policy on those detached targets before an EMA refresh. In tests across SD 3.5-M and Z-Image-Turbo, it reports the best held-out score in 19 of 20 reward-matched settings. The authors also report up to a 44.0% gain over the strongest competing method and lower GPU-hour costs than DiffusionNFT. HF Daily Papers' note

score 5

Categories: Research