arXiv · 1505.00369
Batched bandit problems
Abstract
Motivated by practical applications, chiefly clinical trials, we study the regret achievable for stochastic bandits under the constraint that the employed policy must split trials into a small number of batches. We propose a simple policy, and show that a very small number of batches gives close to minimax optimal regret bounds. As a byproduct, we derive optimal policies with low switching cost for stochastic bandits.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Vianney Perchet, Philippe Rigollet, Sylvain Chassang, Erik Snowberg. 2016-03-29. Batched bandit problems. https://doi.org/10.1214/15-aos1381
Cite the original work for its findings. Save a collection to share your selection of sources.