IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
IDRF fine-tunes few-step masked diffusion generators for reward without rolling out the reference model.
The paper replaces an intractable sequence-level KL penalty with inverse-distillation regularization, which the authors prove can upper-bound the KL to the reference distribution under an optimal auxiliary denoiser. It treats few-step generation as a finite-horizon MDP and optimizes reward with a clipped policy-gradient objective over the student model’s own trajectories. In experiments on DNA, image, and text generation, IDRF reports high reward with up to 32x fewer denoising steps while mitigating reward hacking and preserving sample quality. ArXiv · AI/CL/LG's note
The paper replaces an intractable sequence-level KL penalty with inverse-distillation regularization, which the authors prove can upper-bound the KL to the reference distribution under an optimal auxiliary denoiser. It treats few-step generation as a finite-horizon MDP and optimizes reward with a clipped policy-gradient objective over the student model’s own trajectories. In experiments on DNA, image, and text generation, IDRF reports high reward with up to 32x fewer denoising steps while mitigating reward hacking and preserving sample quality. ArXiv · AI/CL/LG's note
score 5