Closing the Horizon Gap in Policy Optimization for Adversarial MDPs
The paper claims policy optimization can match the horizon dependence of stronger occupancy-measure methods in adversarial episodic MDPs.
Li and Tsuchiya use regularized Q-functions to stabilize local policy updates across all state-action pairs. For known transitions, their algorithm gives a high-probability regret bound of `~O(sqrt(HS(H+A)T))`. For unknown transitions, it gives `~O(HS sqrt(AT))`, matching the best-known bound cited in the abstract. The method is also extended to adversarial linear-mixture MDPs with the same horizon-dependence improvement. Source: ArXiv · AI/CL/LG's note.
Li and Tsuchiya use regularized Q-functions to stabilize local policy updates across all state-action pairs. For known transitions, their algorithm gives a high-probability regret bound of `~O(sqrt(HS(H+A)T))`. For unknown transitions, it gives `~O(HS sqrt(AT))`, matching the best-known bound cited in the abstract. The method is also extended to adversarial linear-mixture MDPs with the same horizon-dependence improvement. Source: ArXiv · AI/CL/LG's note.
score 3