arXiv · 2304.02572
Evaluating Online Bandit Exploration In Large-Scale Recommender System
Abstract
Bandit learning has been an increasingly popular design choice for recommender system. Despite the strong interest in bandit learning from the community, there remains multiple bottlenecks that prevent many bandit learning approaches from productionalization. One major bottleneck is how to test the effectiveness of bandit algorithm with fairness and without data leakage. Different from supervised learning algorithms, bandit learning algorithms emphasize greatly on the data collection process through their explorative nature. Such explorative behavior may induce unfair evaluation in a classic A/B test setting. In this work, we apply upper confidence bound (UCB) to our large scale short video recommender system and present a test framework for the production bandit learning life-cycle with a new set of metrics. Extensive experiment results show that our experiment design is able to fairly evaluate the performance of bandit learning in the recommender system.
Explore related subjects
Keep this discovery
Hongbo Guo, Ruben Naeff, Alex Nikulkov, Zheqing Zhu. 2023-04-05. Evaluating Online Bandit Exploration In Large-Scale Recommender System. https://arxiv.org/abs/2304.02572
Cite the original work for its findings. Save a collection to share your selection of sources.