Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
ERPO moves the drift control from model responses to training queries.
The paper argues that standard Policy-KL regularization limits exploration while removing it leaves policy optimization without explicit drift control. Its proposed Query-KL term bounds shifts in the training-query distribution without applying direct gradient pressure to the response distribution. The method is designed to plug into GRPO, PPO, and REINFORCE-style pipelines without extra forward passes. In tests on six math reasoning benchmarks, it reports stronger accuracy and more stable behavior under high-temperature decoding and long-horizon training. HF Daily Papers' note
The paper argues that standard Policy-KL regularization limits exploration while removing it leaves policy optimization without explicit drift control. Its proposed Query-KL term bounds shifts in the training-query distribution without applying direct gradient pressure to the response distribution. The method is designed to plug into GRPO, PPO, and REINFORCE-style pipelines without extra forward passes. In tests on six math reasoning benchmarks, it reports stronger accuracy and more stable behavior under high-temperature decoding and long-horizon training. HF Daily Papers' note
score 5