arXiv · 2007.08448
Comparator-adaptive Convex Bandits
Abstract
We study bandit convex optimization methods that adapt to the norm of the comparator, a topic that has only been studied before for its full-information counterpart. Specifically, we develop convex bandit algorithms with regret bounds that are small whenever the norm of the comparator is small. We first use techniques from the full-information setting to develop comparator-adaptive algorithms for linear bandits. Then, we extend the ideas to convex bandits with Lipschitz or smooth loss functions, using a new single-point gradient estimator and carefully designed surrogate losses.
Explore related subjects
Keep this discovery
Dirk van der Hoeven, Ashok Cutkosky, Haipeng Luo. 2020-07-16. Comparator-adaptive Convex Bandits. https://arxiv.org/abs/2007.08448
Cite the original work for its findings. Save a collection to share your selection of sources.