Offline Deep Q* Estimation with Diffusion Models
The paper uses conditional diffusion models to estimate the Bellman operator from offline data before learning \(Q^*\).
The authors frame offline RL as solving the optimal Bellman equation when rewards and transitions are not directly known. Their method first learns the reward law and transition kernel, then plugs those estimates into a residual-minimization stage over neural networks. The paper gives nonasymptotic convergence rates for the operator estimate and for the resulting \(Q^*\) estimator, without relying on common completeness assumptions. It also reports numerical experiments showing strong empirical performance. ArXiv · AI/CL/LG's note
The authors frame offline RL as solving the optimal Bellman equation when rewards and transitions are not directly known. Their method first learns the reward law and transition kernel, then plugs those estimates into a residual-minimization stage over neural networks. The paper gives nonasymptotic convergence rates for the operator estimate and for the resulting \(Q^*\) estimator, without relying on common completeness assumptions. It also reports numerical experiments showing strong empirical performance. ArXiv · AI/CL/LG's note
score 5