Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
RGPO uses ground-truth rationales as temporary training scaffolds, then keeps only stronger model-generated answers for unguided reinforcement learning.
The paper frames this as a way to reduce reward sparsity when models cannot find correct reasoning paths on hard problems. It does not require auxiliary examples to match the reinforcement-learning task format. The authors report gains over RLVR baselines in both language-only and vision-language reasoning, with ablations pointing to adaptive rationale guidance as important. HF Daily Papers' note
The paper frames this as a way to reduce reward sparsity when models cannot find correct reasoning paths on hard problems. It does not require auxiliary examples to match the reinforcement-learning task format. The authors report gains over RLVR baselines in both language-only and vision-language reasoning, with ablations pointing to adaptive rationale guidance as important. HF Daily Papers' note
score 4