arXiv · 2007.05456
Improved Analysis of UCRL2 with Empirical Bernstein Inequality
Abstract
We consider the problem of exploration-exploitation in communicating Markov Decision Processes. We provide an analysis of UCRL2 with Empirical Bernstein inequalities (UCRL2B). For any MDP with $S$ states, $A$ actions, $\Gamma \leq S$ next states and diameter $D$, the regret of UCRL2B is bounded as $\widetilde{O}(\sqrt{D\Gamma S A T})$.
Explore related subjects
Keep this discovery
Ronan Fruit, Matteo Pirotta, Alessandro Lazaric. 2020-07-10. Improved Analysis of UCRL2 with Empirical Bernstein Inequality. https://arxiv.org/abs/2007.05456
Cite the original work for its findings. Save a collection to share your selection of sources.