TY - RPRT TI - Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback AU - Zitian Li AU - Wang Chi Cheung PY - 2026 UR - https://arxiv.org/abs/2605.23182 ID - 2605.23182 ER -