Megadose AI progress, ranked and analyzed.

Reward-guided Fine-Tuning of One-Step Generative Models via Wasserstein Gradient Flow

· ArXiv · AI/CL/LG ·
A WGF-based fine-tuning method is proposed for steering one-step generators with rewards, without needing reward gradients.

The paper treats one-step generators through an optimal-transport lens and uses Wasserstein Gradient Flow to control distribution updates. The authors say the method can work with differentiable and non-differentiable rewards while reducing reward hacking and mode collapse. Experiments span 2D synthetic data, CIFAR-10, and ImageNet 256x256, with rewards including JPEG compressibility, class probability, black-and-white output, and CLIP alignment. It reports better reward alignment than baseline methods. ArXiv · AI/CL/LG's note

score 4

Categories: Research