Megadose AI progress, ranked and analyzed.

Achieving an $O(1/N)$ Optimality Gap in Average-Reward Weakly-Coupled MDPs

· ArXiv · AI/CL/LG ·
The paper claims an $O(1/N)$ optimality gap for a general class of weakly-coupled MDPs where earlier general results stopped at $1/\sqrt{N}$.

The authors study average-reward WCMDPs made of $N$ identical arms sharing multiple per-step budget constraints. They identify conditions, analogous to known restless-bandit results, under which a better-than-$1/\sqrt{N}$ gap is possible. Their policy is not priority-based; it is designed to create locally linear mean-field dynamics. ArXiv · AI/CL/LG's note

score 4

Categories: Research