arXiv · 2610.02159
When Do Intrinsic Rewards Lead to Exploration?
Abstract
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Scott W. Viteri, Laura Gomezjurado Gonzalez, Clark Barrett. 2026-10-01. When Do Intrinsic Rewards Lead to Exploration?. https://arxiv.org/abs/2610.02159
Cite the original work for its findings. Save a collection to share your selection of sources.