arXiv · math/0510384
A penalized bandit algorithm
Abstract
We study a two armed-bandit algorithm with penalty. We show the convergence of the algorithm and establish the rate of convergence. For some choices of the parameters, we obtain a central limit theorem in which the limit distribution is characterized as the unique stationary distribution of a discontinuous Markov process.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Damien Lamberton, Gilles Pagès. 2005-10-18. A penalized bandit algorithm. https://arxiv.org/abs/math/0510384
Cite the original work for its findings. Save a collection to share your selection of sources.