Megadose AI progress, ranked and analyzed.

Latent Reward Registers for Diffusion Preference Alignment

· ArXiv · AI/CL/LG ·
The paper’s pitch is a denoising-time reward signal for diffusion alignment, read from noisy latents without changing the generator.

Latent Reward Registers add learnable register tokens to a frozen DiT so terminal preference can be estimated from intermediate latents. That turns a sparse final-sample reward into a dense, differentiable signal across the generation trajectory. The authors use it for RG-OPD training and RGS inference-time steering, reporting up to 33x fewer GPU hours than online RL baselines and state-of-the-art results among training-free methods. ArXiv · AI/CL/LG's note

score 5

Categories: Research