TY - RPRT TI - Non-Asymptotic Best Policy Identification Guarantees in Online Reinforcement Learning AU - Joseph Lazzaro AU - Alessio Russo AU - Aldo Pacchiano PY - 2026 UR - https://arxiv.org/abs/2607.17201 ID - 2607.17201 ER -