Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
UMM-Reflection trains one multimodal model to critique and revise its own image generations across full repair loops.
The paper says supervised reflection traces help initialize the system, but do not find the strongest repair paths. Its interleaved RL method scores complete sibling trajectories from the same starting image, then updates both the reflection text and image revisions with one trajectory-level advantage. On BAGEL, it reports a 12.05-point GenEval gain over SFT, with transfer gains on WISE, OneIG-Bench, and T2I-CompBench++. HF Daily Papers' note
The paper says supervised reflection traces help initialize the system, but do not find the strongest repair paths. Its interleaved RL method scores complete sibling trajectories from the same starting image, then updates both the reflection text and image revisions with one trajectory-level advantage. On BAGEL, it reports a 12.05-point GenEval gain over SFT, with transfer gains on WISE, OneIG-Bench, and T2I-CompBench++. HF Daily Papers' note
score 5