Megadose Built for builders and researchers.

Closing the Horizon Gap in Policy Optimization for Adversarial MDPs

· ArXiv · AI/CL/LG ·
The paper claims policy optimization can match the horizon dependence of stronger occupancy-measure methods in adversarial episodic MDPs.

Li and Tsuchiya use regularized Q-functions to stabilize local policy updates across all state-action pairs. For known transitions, their algorithm gives a high-probability regret bound of `~O(sqrt(HS(H+A)T))`. For unknown transitions, it gives `~O(HS sqrt(AT))`, matching the best-known bound cited in the abstract. The method is also extended to adversarial linear-mixture MDPs with the same horizon-dependence improvement. Source: ArXiv · AI/CL/LG's note.

score 3

Categories: Research