TY - RPRT TI - Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time Oracles AU - Bhrij Patel AU - Wesley A. Suttle AU - Alec Koppel AU - Vaneet Aggarwal AU - Brian M. Sadler AU - Amrit Singh Bedi AU - Dinesh Manocha PY - 2024 UR - https://arxiv.org/abs/2403.11925 ID - 2403.11925 ER -