arXiv · 2301.11442
Communication-Efficient Collaborative Regret Minimization in Multi-Armed Bandits
Abstract
In this paper, we study the collaborative learning model, which concerns the tradeoff between parallelism and communication overhead in multi-agent multi-armed bandits. For regret minimization in multi-armed bandits, we present the first set of tradeoffs between the number of rounds of communication among the agents and the regret of the collaborative learning process.
Explore related subjects
Keep this discovery
Nikolai Karpov, Qin Zhang. 2023-01-26. Communication-Efficient Collaborative Regret Minimization in Multi-Armed Bandits. https://arxiv.org/abs/2301.11442
Cite the original work for its findings. Save a collection to share your selection of sources.