Megadose AI progress, ranked and analyzed.

Q-based Variational Inverse Reinforcement Learning

· ArXiv · AI/CL/LG ·
QVIRL estimates uncertainty over learned rewards by training mainly on optimal Q-values.

The paper presents a Bayesian inverse reinforcement learning method meant to infer human preferences from expert behavior. Its claimed advantage is combining scalability with posterior uncertainty, which the authors tie to safety-critical use and active learning. They report strong apprenticeship-learning results on gridworlds, Lunar Lander, Highway Environment, and two ATARI games. The authors also say it is the first Bayesian IRL method shown training from raw pixel observations. ArXiv · AI/CL/LG's note

score 4

Categories: Research