arXiv · 2106.14308
Concentration of Contractive Stochastic Approximation and Reinforcement Learning
Abstract
Using a martingale concentration inequality, concentration bounds `from time $n_0$ on' are derived for stochastic approximation algorithms with contractive maps and both martingale difference and Markov noises. These are applied to reinforcement learning algorithms, in particular to asynchronous Q-learning and TD(0).
Explore related subjects
Keep this discovery
Siddharth Chandak, Vivek S. Borkar, Parth Dodhia. 2021-06-27. Concentration of Contractive Stochastic Approximation and Reinforcement Learning. https://arxiv.org/abs/2106.14308
Cite the original work for its findings. Save a collection to share your selection of sources.