RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction
The paper says generative reward models work better for RL when their rankings are converted into relative rewards, not forced into scalar scores.
RRC builds rewards from preference orderings using self-competitive comparisons among sampled answers and anchor-guided comparisons against a small reference set. The authors say this addresses the mismatch between ranking-based reward models and standard RL training signals. In experiments on open-ended chat and reasoning benchmarks, RRC produced consistent gains over existing reward construction methods. ArXiv · AI/CL/LG's note
RRC builds rewards from preference orderings using self-competitive comparisons among sampled answers and anchor-guided comparisons against a small reference set. The authors say this addresses the mismatch between ranking-based reward models and standard RL training signals. In experiments on open-ended chat and reasoning benchmarks, RRC produced consistent gains over existing reward construction methods. ArXiv · AI/CL/LG's note
score 5