Role-Adaptive Policy Optimization for Offline Reinforcement Learning
RAPO splits policy tuning by job: one actor for bootstrapping value targets, another for execution.
The paper argues that offline RL methods such as TD3+BC over-couple those roles through a shared policy. RAPO adapts update coefficients separately, penalizing bootstrap changes that disturb target values while using a local improvement surrogate for the execution actor. For IQL, it leaves value learning intact and adapts only the inverse temperature used in advantage-weighted policy extraction. The authors report gains on D4RL locomotion and AntMaze, with the larger improvements coming from the TD3+BC version. ArXiv · AI/CL/LG's note
The paper argues that offline RL methods such as TD3+BC over-couple those roles through a shared policy. RAPO adapts update coefficients separately, penalizing bootstrap changes that disturb target values while using a local improvement surrogate for the execution actor. For IQL, it leaves value learning intact and adapts only the inverse temperature used in advantage-weighted policy extraction. The authors report gains on D4RL locomotion and AntMaze, with the larger improvements coming from the TD3+BC version. ArXiv · AI/CL/LG's note
score 4