Megadose AI progress, ranked and analyzed.

Evidence-RL: Towards Evidence-intensive Visual Reasoning

· HF Daily Papers ·
The paper trains VLMs to prove their answers depend on the right image region.

Evidence-RL introduces Counterfactual Evidence Disentanglement, which masks an object-centered evidence region and checks whether model support drops more than it does for matched non-evidence regions. That signal is folded into GRPO with answer correctness, rewarding answers that use the evidence path instead of shortcuts or irrelevant context. The authors say it needs only weak object proposals, no question-specific evidence labels, and adds no inference-time cost. They report gains over prior RL post-training methods across nine benchmarks and four backbones. HF Daily Papers' note

score 5

Categories: Research