arXiv · 2009.12511
Near-Optimal MNL Bandits Under Risk Criteria
Abstract
We study MNL bandits, which is a variant of the traditional multi-armed bandit problem, under risk criteria. Unlike the ordinary expected revenue, risk criteria are more general goals widely used in industries and bussiness. We design algorithms for a broad class of risk criteria, including but not limited to the well-known conditional value-at-risk, Sharpe ratio and entropy risk, and prove that they suffer a near-optimal regret. As a complement, we also conduct experiments with both synthetic and real data to show the empirical performance of our proposed algorithms.
Explore related subjects
Keep this discovery
Guangyu Xi, Chao Tao, Yuan Zhou. 2020-09-26. Near-Optimal MNL Bandits Under Risk Criteria. https://arxiv.org/abs/2009.12511
Cite the original work for its findings. Save a collection to share your selection of sources.