arXiv · 2110.05724
Query-Reward Tradeoffs in Multi-Armed Bandits
Abstract
We consider a stochastic multi-armed bandit setting where reward must be actively queried for it to be observed. We provide tight lower and upper problem-dependent guarantees on both the regret and the number of queries. Interestingly, we prove that there is a fundamental difference between problems with a unique and multiple optimal arms, unlike in the standard multi-armed bandit problem. We also present a new, simple, UCB-style sampling concept, and show that it naturally adapts to the number of optimal arms and achieves tight regret and querying bounds.
Explore related subjects
Keep this discovery
Nadav Merlis, Yonathan Efroni, Shie Mannor. 2021-10-12. Query-Reward Tradeoffs in Multi-Armed Bandits. https://arxiv.org/abs/2110.05724
Cite the original work for its findings. Save a collection to share your selection of sources.