TY - RPRT TI - Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs AU - Jiafan He AU - Dongruo Zhou AU - Quanquan Gu PY - 2022 UR - https://arxiv.org/abs/2010.00587 ID - 2010.00587 ER -