Megadose AI progress, ranked and analyzed.

Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging

· ArXiv · AI/CL/LG ·
A small POMDP construction can make logged data exponentially useless even when coverage looks good.

The paper shows this under a history-dependent logger: for horizon \(H \ge 3\), off-policy evaluation to accuracy \(1/8\) needs \(\Theta((3/2)^H \log(1/\delta))\) logged episodes. The construction keeps action coverage, belief coverage, and outcome-revealing conditions bounded independently of \(H\), but a reset hides the transition that fixes the target policy’s value. The author also gives an exact statistical characterization, a matching optimal estimator, and a directed two-lane gridworld realization. Source: ArXiv · AI/CL/LG's note.

score 4

Categories: Research