OR Else: A Differentiable Trust Region for Policy Optimization
Output Reset improved PPO’s training reward score here, but did not lift GRPO at the tested group size.
The paper swaps clipped policy terms for a smooth one-sided OR squared-margin loss in PPO and GRPO variants. On Llama-3.2-1B-Instruct with Anthropic hh-rlhf, PPO-OR finished 0.305 higher than PPO-clip on mean training-time reward-model score, with more seed spread. GRPO-OR showed steadier diagnostics than GRPO, including near-zero terminal OR residual, but no higher mean reward score at group size 2. The authors stress these are training-time reward-model results, not held-out human-preference performance. ArXiv · AI/CL/LG's note
The paper swaps clipped policy terms for a smooth one-sided OR squared-margin loss in PPO and GRPO variants. On Llama-3.2-1B-Instruct with Anthropic hh-rlhf, PPO-OR finished 0.305 higher than PPO-clip on mean training-time reward-model score, with more seed spread. GRPO-OR showed steadier diagnostics than GRPO, including near-zero terminal OR residual, but no higher mean reward score at group size 2. The authors stress these are training-time reward-model results, not held-out human-preference performance. ArXiv · AI/CL/LG's note
score 5