TY - RPRT TI - Regret Bounds for Information-Directed Reinforcement Learning AU - Botao Hao AU - Tor Lattimore PY - 2022 UR - https://arxiv.org/abs/2206.04640 ID - 2206.04640 ER -