Online Inference in Distributional Temporal-Difference Learning
The paper gives bootstrap-valid uncertainty estimates for return-distribution TD learning from one Markov trajectory.
Peng and Zhang prove a root-\(T\) Gaussian limit for the Polyak-Ruppert averaged estimator in nonparametric distributional temporal-difference learning. They show the bootstrap matches that limit conditionally on the observed path, supporting inference for smooth functionals such as variance, CVaR, expected shortfall, and expectiles. For quantiles and other nonsmooth targets, they develop a local asymptotic theory around CDF thresholds. ArXiv · AI/CL/LG's note
Peng and Zhang prove a root-\(T\) Gaussian limit for the Polyak-Ruppert averaged estimator in nonparametric distributional temporal-difference learning. They show the bootstrap matches that limit conditionally on the observed path, supporting inference for smooth functionals such as variance, CVaR, expected shortfall, and expectiles. For quantiles and other nonsmooth targets, they develop a local asymptotic theory around CDF thresholds. ArXiv · AI/CL/LG's note
score 4