MEND: RL For Flow Models via Proximal Velocity Matching
MEND trains flow models by accepting only reward-gradient moves whose capped gain justifies the displacement.
The paper frames this as a no-KL, no-reference alternative to reward post-training methods that need many updates or move samples indiscriminately. In its reported runs, MEND beats Flow-GRPO on five of six evaluators in 100 updates, compared with about 4,000 for Flow-GRPO. Under an equal-budget setup, it also outperforms ReFL and DiffusionNFT across four training rewards, including a PickScore of 24.03. The authors say it can be applied to any flow backbone with a differentiable reward. HF Daily Papers' note
The paper frames this as a no-KL, no-reference alternative to reward post-training methods that need many updates or move samples indiscriminately. In its reported runs, MEND beats Flow-GRPO on five of six evaluators in 100 updates, compared with about 4,000 for Flow-GRPO. Under an equal-budget setup, it also outperforms ReFL and DiffusionNFT across four training rewards, including a PickScore of 24.03. The authors say it can be applied to any flow backbone with a differentiable reward. HF Daily Papers' note
score 5