When Do Intrinsic Rewards Lead to Exploration?
Some intrinsic-reward objectives can optimize cleanly and still fail to gather the most useful information.
The paper proposes judging exploration by counterfactual information: whether an agent’s history can stand in for what it would have learned under other policies. In one simple environment, count-based, prediction-error, empowerment, and information-gain objectives all have maximizing policies that are Pareto-suboptimal by that test. The authors then identify conditions where existing intrinsic rewards do lead to optimal exploration, and define an objective that increases whenever exploration improves under their criterion. ArXiv · AI/CL/LG's note
The paper proposes judging exploration by counterfactual information: whether an agent’s history can stand in for what it would have learned under other policies. In one simple environment, count-based, prediction-error, empowerment, and information-gain objectives all have maximizing policies that are Pareto-suboptimal by that test. The authors then identify conditions where existing intrinsic rewards do lead to optimal exploration, and define an objective that increases whenever exploration improves under their criterion. ArXiv · AI/CL/LG's note
score 4