TY - RPRT TI - Optimistically Optimistic Exploration for Provably Efficient Infinite-Horizon Reinforcement and Imitation Learning AU - Antoine Moulin AU - Gergely Neu AU - Luca Viano PY - 2026 UR - https://arxiv.org/abs/2502.13900 ID - 2502.13900 ER -