Verifier Errors in RLVR: Reward Hacking, Limits of Feedback, and Selective Control
Imperfect verifiers can make RLVR optimize for accepted wrong answers instead of correctness.
The paper characterizes when verifier reward can rise while actual correctness falls. It argues that the signals normally visible during RLVR are not enough to detect accepted errors, identify them, or reliably reduce them without also losing correct responses. The authors propose adding audit feedback about correctness, producing a correction that lowers accepted errors and raises correct responses at the current policy when strong enough. Experiments on bandits and a language model support that selective-control approach. ArXiv · AI/CL/LG's note
The paper characterizes when verifier reward can rise while actual correctness falls. It argues that the signals normally visible during RLVR are not enough to detect accepted errors, identify them, or reliably reduce them without also losing correct responses. The authors propose adding audit feedback about correctness, producing a correction that lowers accepted errors and raises correct responses at the current policy when strong enough. Experiments on bandits and a language model support that selective-control approach. ArXiv · AI/CL/LG's note
score 6