iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
The paper says diffusion reward tuning can damage diversity, and argues iADD improves that tradeoff.
The authors challenge prior claims that updating only later diffusion timesteps protects diversity. They propose incremental Feynman-Kac training as the basis for iADD. In experiments across three tasks, they report gains in both alignment and diversity against related diffusion policy optimization methods. HF Daily Papers' note
The authors challenge prior claims that updating only later diffusion timesteps protects diversity. They propose incremental Feynman-Kac training as the basis for iADD. In experiments across three tasks, they report gains in both alignment and diversity against related diffusion policy optimization methods. HF Daily Papers' note
score 4