Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
The paper argues that rival diffusion-RL training losses are the same path-space objective seen through different variance choices.
It derives a trajectory-space policy-gradient estimator from a regularized diffusion-RL objective using importance sampling between sampling SDEs. The authors say Flow-GRPO-style reverse-trajectory updates and forward-matching methods such as AWM and DiffusionNFT fall out of that same derivation. The reported split between those families is framed as variance reduction, not a different reinforcement-learning principle. They also propose a KDE value-gradient estimator and bounded weighting choices, with experiments on SD3.5-M and Qwen-Image beating prior diffusion-RL baselines. ArXiv · AI/CL/LG's note
It derives a trajectory-space policy-gradient estimator from a regularized diffusion-RL objective using importance sampling between sampling SDEs. The authors say Flow-GRPO-style reverse-trajectory updates and forward-matching methods such as AWM and DiffusionNFT fall out of that same derivation. The reported split between those families is framed as variance reduction, not a different reinforcement-learning principle. They also propose a KDE value-gradient estimator and bounded weighting choices, with experiments on SD3.5-M and Qwen-Image beating prior diffusion-RL baselines. ArXiv · AI/CL/LG's note
score 5