Verifier-Induced Support Reshaping in On-Policy Optimization
RLVR can raise the score in front of it while making later-successful behaviors harder to find.
The paper names this “verifier-induced support reshaping”: successful trajectories for a future objective can become too rare to sample within a fixed rollout budget. In Qwen3-8B-Base on IFEval, Math-RLVR lifted pass@1 by 6.5 points while best@32 fell by 9.8 points. The authors trace much of the shift to the first few response tokens, where RLVR tends to rerank openings already present in the base policy. Tested constraints and routing methods only partly preserved cross-task support. HF Daily Papers' note
The paper names this “verifier-induced support reshaping”: successful trajectories for a future objective can become too rare to sample within a fixed rollout budget. In Qwen3-8B-Base on IFEval, Math-RLVR lifted pass@1 by 6.5 points while best@32 fell by 9.8 points. The authors trace much of the shift to the first few response tokens, where RLVR tends to rerank openings already present in the base policy. Tested constraints and routing methods only partly preserved cross-task support. HF Daily Papers' note
score 5