arXiv · 2610.10128
m-Set Adversarial Bandits with Winner Feedback
Abstract
We show upper and lower bounds on the regret of $m$-set adversarial bandits for different utilities (winner reward or sum of rewards) and feedback models (winner index, winner reward, sum of rewards, and their combinations). By comparing to standard bounds for combinatorial and MNL bandits, our results reveal how subtle changes in the setting can have a dramatic impact on the learning rates. Our main technical contributions are the information-theoretic lower bounds on the regret. Experiments on synthetic data confirm our theoretical analyses.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nicolò Cesa-Bianchi, Matteo Papini. 2026-10-07. m-Set Adversarial Bandits with Winner Feedback. https://arxiv.org/abs/2610.10128
Cite the original work for its findings. Save a collection to share your selection of sources.