Megadose AI progress, ranked and analyzed.

OR Else: A Differentiable Trust Region for Policy Optimization

· ArXiv · AI/CL/LG ·
Output Reset improved PPO’s training reward score here, but did not lift GRPO at the tested group size.

The paper swaps clipped policy terms for a smooth one-sided OR squared-margin loss in PPO and GRPO variants. On Llama-3.2-1B-Instruct with Anthropic hh-rlhf, PPO-OR finished 0.305 higher than PPO-clip on mean training-time reward-model score, with more seed spread. GRPO-OR showed steadier diagnostics than GRPO, including near-zero terminal OR residual, but no higher mean reward score at group size 2. The authors stress these are training-time reward-model results, not held-out human-preference performance. ArXiv · AI/CL/LG's note

score 5

Categories: Research