Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
RWTD trains one-step image generators from samples and scalar rewards, without likelihoods, denoising paths, or differentiable rewards.
The method builds an adaptive target from both the current model and the pretrained reference, aiming to improve reward alignment while retaining prior capability. It uses feature-space optimal transport and fixed-point regression to realize that target. In experiments, it raises the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80. The authors also report preference-alignment tests with cross-reward generalization and preserved compositional behavior. HF Daily Papers' note
The method builds an adaptive target from both the current model and the pretrained reference, aiming to improve reward alignment while retaining prior capability. It uses feature-space optimal transport and fixed-point regression to realize that target. In experiments, it raises the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80. The authors also report preference-alignment tests with cross-reward generalization and preserved compositional behavior. HF Daily Papers' note
score 5