Megadose Built for builders and researchers.

Policy as Data: Replay-Based Policy Dual Averaging via Advantage Regression

· ArXiv · AI/CL/LG ·
RDA2C treats replay as accumulated policy-improvement data, not just off-policy residue.

The paper introduces Regularized Dual Averaging Actor Critic, which stores state-action samples with critic-estimated advantage labels and fits a dual score model over the aggregated replay set. The current policy is then derived from that accumulated score model through an entropy mirror map. The authors give a finite-time value-gap decomposition covering dual averaging, stale replay fitting, critic bias, buffer variance, and replay coverage. In reported tests, GAE-labeled RDA2C beats PPO on six of eight MuJoCo tasks and eight of twelve Atari games, and twin-Q labels let it match SAC under matched batch size and update frequency. ArXiv · AI/CL/LG's note

score 4

Categories: Research