SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
The model uses its own Yes/No self-check as the reinforcement signal for vision-language reasoning.
SVR-R1 has the same model propose an answer, judge it, and retry when the verdict is No. A Yes verdict, or a turn limit, locks the answer for outcome-based reward under GRPO. The authors say it needs no external supervision or auxiliary critic, and beats strong GRPO baselines on vision-language reasoning benchmarks. They also report that verification turns fall during training while test accuracy rises. HF Daily Papers' note
SVR-R1 has the same model propose an answer, judge it, and retry when the verdict is No. A Yes verdict, or a turn limit, locks the answer for outcome-based reward under GRPO. The authors say it needs no external supervision or auxiliary critic, and beats strong GRPO baselines on vision-language reasoning benchmarks. They also report that verification turns fall during training while test accuracy rises. HF Daily Papers' note
score 6