arXiv · 2211.17154
On Regret-optimal Cooperative Nonstochastic Multi-armed Bandits
Abstract
We consider the nonstochastic multi-agent multi-armed bandit problem with agents collaborating via a communication network with delays. We show a lower bound for individual regret of all agents. We show that with suitable regularizers and communication protocols, a collaborative multi-agent \emph{follow-the-regularized-leader} (FTRL) algorithm has an individual regret upper bound that matches the lower bound up to a constant factor when the number of arms is large enough relative to degrees of agents in the communication graph. We also show that an FTRL algorithm with a suitable regularizer is regret optimal with respect to the scaling with the edge-delay parameter. We present numerical experiments validating our theoretical results and demonstrate cases when our algorithms outperform previously proposed algorithms.
Explore related subjects
Keep this discovery
Jialin Yi, Milan Vojnović. 2022-11-30. On Regret-optimal Cooperative Nonstochastic Multi-armed Bandits. https://doi.org/10.5555/3545946.3598780
Cite the original work for its findings. Save a collection to share your selection of sources.