Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
The paper proposes EVR, a reward model meant to make multi-reference image edits more visually consistent.
EVR splits judging into separate visual criteria, then has an MLLM generate candidate claims and a verifier accept or reject them against visual evidence. The authors say this avoids relying on either long hallucination-prone reasoning or shallow short judgments. Used with a scalable data pipeline, the reward lets off-the-shelf editors be RL fine-tuned without changing their architecture. Experiments report gains over base Qwen-Image-Edit, with consistency and harmony matching or surpassing NanoBanana. HF Daily Papers' note
EVR splits judging into separate visual criteria, then has an MLLM generate candidate claims and a verifier accept or reject them against visual evidence. The authors say this avoids relying on either long hallucination-prone reasoning or shallow short judgments. Used with a scalable data pipeline, the reward lets off-the-shelf editors be RL fine-tuned without changing their architecture. Experiments report gains over base Qwen-Image-Edit, with consistency and harmony matching or surpassing NanoBanana. HF Daily Papers' note
score 4