Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
UMM-Reflection trains one multimodal model to critique and revise its own image generations across full repair loops.
The paper says supervised reflection traces help start the process but miss better repair paths. Its method applies reinforcement learning over complete reflection trajectories, updating both the reflection text and the flow-based image revisions. The authors report a 12.05-point GenEval gain over SFT on BAGEL, with transfer gains on WISE, OneIG-Bench, and T2I-CompBench++ without training on them. ArXiv · AI/CL/LG's note
The paper says supervised reflection traces help start the process but miss better repair paths. Its method applies reinforcement learning over complete reflection trajectories, updating both the reflection text and the flow-based image revisions. The authors report a 12.05-point GenEval gain over SFT on BAGEL, with transfer gains on WISE, OneIG-Bench, and T2I-CompBench++ without training on them. ArXiv · AI/CL/LG's note
score 5