Verifiable Visual Rewards Transfer from Synthetic Scenes to Natural Prompts
The paper claims deterministic synthetic-scene checks can train image models to follow harder visual instructions.
VVR generates geometric-object tasks with paired prompts and programmatic verifiers, avoiding reward signals from less reliable detectors or vision-language models. The authors release VVRBench with 10,000 tasks across 32 constraint types, plus a 720-task challenge set. GPT-Image-2.5 solves 21.4% of the challenge benchmark in their evaluation. Reinforcement learning with VVR rewards lifts Stable Diffusion 3.5 Medium from 2.8% to 28.3% on VVRBench, with reported gains carrying to out-of-domain benchmarks.
ArXiv · AI/CL/LG's note
VVR generates geometric-object tasks with paired prompts and programmatic verifiers, avoiding reward signals from less reliable detectors or vision-language models. The authors release VVRBench with 10,000 tasks across 32 constraint types, plus a 720-task challenge set. GPT-Image-2.5 solves 21.4% of the challenge benchmark in their evaluation. Reinforcement learning with VVR rewards lifts Stable Diffusion 3.5 Medium from 2.8% to 28.3% on VVRBench, with reported gains carrying to out-of-domain benchmarks.
ArXiv · AI/CL/LG's note
score 6