arXiv · 2007.08703
Bandits for BMO Functions
Abstract
We study the bandit problem where the underlying expected reward is a Bounded Mean Oscillation (BMO) function. BMO functions are allowed to be discontinuous and unbounded, and are useful in modeling signals with infinities in the do-main. We develop a toolset for BMO bandits, and provide an algorithm that can achieve poly-log $\delta$-regret -- a regret measured against an arm that is optimal after removing a $\delta$-sized portion of the arm space.
Explore related subjects
Keep this discovery
Tianyu Wang, Cynthia Rudin. 2020-07-17. Bandits for BMO Functions. https://arxiv.org/abs/2007.08703
Cite the original work for its findings. Save a collection to share your selection of sources.