arXiv · 1311.0468
Thompson Sampling for Online Learning with Linear Experts
Abstract
In this note, we present a version of the Thompson sampling algorithm for the problem of online linear generalization with full information (i.e., the experts setting), studied by Kalai and Vempala, 2005. The algorithm uses a Gaussian prior and time-varying Gaussian likelihoods, and we show that it essentially reduces to Kalai and Vempala's Follow-the-Perturbed-Leader strategy, with exponentially distributed noise replaced by Gaussian noise. This implies sqrt(T) regret bounds for Thompson sampling (with time-varying likelihood) for online learning with full information.
Explore related subjects
Keep this discovery
Aditya Gopalan. 2013-11-03. Thompson Sampling for Online Learning with Linear Experts. https://arxiv.org/abs/1311.0468
Cite the original work for its findings. Save a collection to share your selection of sources.