Megadose AI progress, ranked and analyzed.

Group Adaptive Clipping Policy Optimization

· HF Daily Papers ·
The paper argues fixed GRPO clipping suppresses some of the most useful learning signals on harder problems.

GAPO changes only the clipping threshold, giving rollouts with larger advantage more update headroom. The authors frame this through a reverse-KL trust-region view and keep the standard PPO/GSPO surrogate intact. In tests on Qwen and Llama models, it improves Pass@1 and Pass@k over fixed clipping and advantage-shaping baselines on lower-pass-rate math and coding benchmarks. HF Daily Papers' note

score 5

Categories: Research