FERPO: Forward Entropy-Regularized Policy Optimization
FERPO updates policies from critic values without taking action gradients through the critic.
The paper argues that value critics can predict returns well while still giving unreliable action derivatives. FERPO instead builds an entropy- and KL-regularized target action distribution, then fits the actor to it with a forward-KL objective estimated by self-normalized importance sampling. The authors say this helps cover multiple high-value modes, supporting exploration, while KL regularization keeps the sampling weights controlled. Experiments on MuJoCo Playground and ManiSkill show competitive results, sample-efficiency gains, and faster actor updates than REPPO. ArXiv · AI/CL/LG's note
The paper argues that value critics can predict returns well while still giving unreliable action derivatives. FERPO instead builds an entropy- and KL-regularized target action distribution, then fits the actor to it with a forward-KL objective estimated by self-normalized importance sampling. The authors say this helps cover multiple high-value modes, supporting exploration, while KL regularization keeps the sampling weights controlled. Experiments on MuJoCo Playground and ManiSkill show competitive results, sample-efficiency gains, and faster actor updates than REPPO. ArXiv · AI/CL/LG's note
score 4