Megadose AI progress, ranked and analyzed.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

· ArXiv · AI/CL/LG ·
On-policy expert queries help most when the learner can model the expert’s value function but not the expert’s full policy.

The paper introduces OVI, an interactive imitation-learning algorithm built around value estimation rather than direct action-distribution matching. Its central claim is that interaction lowers the representational burden: value-function realizability can be enough where policy realizability is not. The authors also give a negative result that offline imitation learning still has to scale with the complexity of the expert policy class under those weaker assumptions. In experiments, OVI beats BC, DAgger, and offline value-based methods, especially when the learner network is much less expressive than the expert’s. ArXiv · AI/CL/LG's note

score 4

Categories: Research