Megadose Built for builders and researchers.

On KL-Regularized Policy Optimization

· HF Daily Papers ·
KLPO targets stale-rollout LLM agent training without importance weights, critics, or grouped responses.

The paper proposes anchoring the KL regularizer at the sampler, so updates are fit on the sampler’s own trajectories through a log-ratio condition.
It claims a closed-form Gibbs improvement step and replaces the log-partition term with a sampler mean plus sampler-to-trainer KL divergence.
For token-level mirror descent targets, the gradient can be computed from terminal returns, including with stochastic tool outputs.
The result is a one-rollout-per-prompt method that the paper says contains SPPO, GPO, REBEL, and BPO as special cases.
Source: HF Daily Papers' note

score 5

Categories: Research