TY - RPRT TI - Nearly Minimax Optimal Regret for Learning Infinite-horizon Average-reward MDPs with Linear Function Approximation AU - Yue Wu AU - Dongruo Zhou AU - Quanquan Gu PY - 2022 UR - https://arxiv.org/abs/2102.07301 ID - 2102.07301 ER -