arXiv · 1709.03570
A KL-LUCB Bandit Algorithm for Large-Scale Crowdsourcing
Abstract
This paper focuses on best-arm identification in multi-armed bandits with bounded rewards. We develop an algorithm that is a fusion of lil-UCB and KL-LUCB, offering the best qualities of the two algorithms in one method. This is achieved by proving a novel anytime confidence bound for the mean of bounded distributions, which is the analogue of the LIL-type bounds recently developed for sub-Gaussian distributions. We corroborate our theoretical results with numerical experiments based on the New Yorker Cartoon Caption Contest.
Explore related subjects
Keep this discovery
Bob Mankoff, Robert Nowak, Ervin Tanczos. 2017-09-11. A KL-LUCB Bandit Algorithm for Large-Scale Crowdsourcing. https://arxiv.org/abs/1709.03570
Cite the original work for its findings. Save a collection to share your selection of sources.