TY - RPRT TI - Unified theory of upper confidence bound policies for bandit problems targeting total reward, maximal reward, and more AU - Nobuaki Kikkawa AU - Hiroshi Ohno PY - 2024 UR - https://arxiv.org/abs/2411.00339 ID - 2411.00339 ER -