arXiv · 1910.02140
Discounted Reinforcement Learning Is Not an Optimization Problem
Abstract
Discounted reinforcement learning is fundamentally incompatible with function approximation for control in continuing tasks. It is not an optimization problem in its usual formulation, so when using function approximation there is no optimal policy. We substantiate these claims, then go on to address some misconceptions about discounting and its connection to the average reward formulation. We encourage researchers to adopt rigorous optimization approaches, such as maximizing average reward, for reinforcement learning in continuing tasks.
Explore related subjects
Keep this discovery
Abhishek Naik, Roshan Shariff, Niko Yasui, Hengshuai Yao, Richard S. Sutton. 2019-10-04. Discounted Reinforcement Learning Is Not an Optimization Problem. https://arxiv.org/abs/1910.02140
Cite the original work for its findings. Save a collection to share your selection of sources.