arXiv · 2311.08376
Ensemble sampling for linear bandits: small ensembles suffice
Abstract
We provide the first useful and rigorous analysis of ensemble sampling for the stochastic linear bandit setting. In particular, we show that, under standard assumptions, for a $d$-dimensional stochastic linear bandit with an interaction horizon $T$, ensemble sampling with an ensemble of size of order $d \log T$ incurs regret at most of the order $(d \log T)^{5/2} \sqrt{T}$. Ours is the first result in any structured setting not to require the size of the ensemble to scale linearly with $T$ -- which defeats the purpose of ensemble sampling -- while obtaining near $\smash{\sqrt{T}}$ order regret. Our result is also the first to allow for infinite action sets.
Explore related subjects
Keep this discovery
David Janz, Alexander E. Litvak, Csaba Szepesvári. 2023-11-14. Ensemble sampling for linear bandits: small ensembles suffice. https://arxiv.org/abs/2311.08376
Cite the original work for its findings. Save a collection to share your selection of sources.