Provable Benefits of Regularization: Fast Rates for Adversarial Imitation Learning
The paper claims jointly regularized AIL can reach fast finite-sample rates in both demonstrations and online interaction.
The authors introduce Dually Regularized AIL, combining KL policy regularization with a quadratic reward penalty. For fixed regularization parameters, they prove a \(\widetilde{O}(1/K + 1/N)\) bound on the regularized imitation gap, where \(K\) is online episodes and \(N\) is expert trajectories. They say this gives \(\widetilde{O}(1/\epsilon)\) sample complexity for both expert data and learner interaction, including stochastic experts. ArXiv · AI/CL/LG's note
The authors introduce Dually Regularized AIL, combining KL policy regularization with a quadratic reward penalty. For fixed regularization parameters, they prove a \(\widetilde{O}(1/K + 1/N)\) bound on the regularized imitation gap, where \(K\) is online episodes and \(N\) is expert trajectories. They say this gives \(\widetilde{O}(1/\epsilon)\) sample complexity for both expert data and learner interaction, including stochastic experts. ArXiv · AI/CL/LG's note
score 4