Megadose Built for builders and researchers.

Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

· HF Daily Papers ·
ERPO moves the drift control from model responses to training queries.

The paper argues that standard Policy-KL regularization limits exploration while removing it leaves policy optimization without explicit drift control. Its proposed Query-KL term bounds shifts in the training-query distribution without applying direct gradient pressure to the response distribution. The method is designed to plug into GRPO, PPO, and REINFORCE-style pipelines without extra forward passes. In tests on six math reasoning benchmarks, it reports stronger accuracy and more stable behavior under high-temperature decoding and long-horizon training. HF Daily Papers' note

score 5

Categories: Research