arXiv · 2011.01075
A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting
Abstract
Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in the finite-horizon case. In this note we show that once adapted to the discounted setting, the construction can be simplified to a 2-state MDP with 1-dimensional features, such that learning is impossible even with an infinite amount of data.
Explore related subjects
Keep this discovery
Philip Amortila, Nan Jiang, Tengyang Xie. 2020-11-02. A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting. https://arxiv.org/abs/2011.01075
Cite the original work for its findings. Save a collection to share your selection of sources.