Megadose Built for builders and researchers.

FERPO: Forward Entropy-Regularized Policy Optimization

· ArXiv · AI/CL/LG ·
FERPO updates policies from critic values without taking action gradients through the critic.

The paper argues that value critics can predict returns well while still giving unreliable action derivatives. FERPO instead builds an entropy- and KL-regularized target action distribution, then fits the actor to it with a forward-KL objective estimated by self-normalized importance sampling. The authors say this helps cover multiple high-value modes, supporting exploration, while KL regularization keeps the sampling weights controlled. Experiments on MuJoCo Playground and ManiSkill show competitive results, sample-efficiency gains, and faster actor updates than REPPO. ArXiv · AI/CL/LG's note

score 4

Categories: Research