EasyPPO: Stabilizing the Critic Is Key
The paper argues PPO’s critic is what often destabilizes LLM reinforcement learning, and proposes three targeted fixes.
EasyPPO keeps critic training on both completed and truncated rollouts while filtering overlong samples only for the actor. It also normalizes critic loss by per-prompt return noise and uses smaller critic mini-batches to limit outlier damage during clipping. The authors report stable full-horizon training across coding, math, and multi-turn search benchmarks, with best validation gains over vanilla PPO of 14.89%, 2.28%, and 9.47%. Source: HF Daily Papers' note.
EasyPPO keeps critic training on both completed and truncated rollouts while filtering overlong samples only for the actor. It also normalizes critic loss by per-prompt return noise and uses smaller critic mini-batches to limit outlier damage during clipping. The authors report stable full-horizon training across coding, math, and multi-turn search benchmarks, with best validation gains over vanilla PPO of 14.89%, 2.28%, and 9.47%. Source: HF Daily Papers' note.
score 5