TY - RPRT TI - Sample-efficient Learning of Infinite-horizon Average-reward MDPs with General Function Approximation AU - Jianliang He AU - Han Zhong AU - Zhuoran Yang PY - 2024 UR - https://arxiv.org/abs/2404.12648 ID - 2404.12648 ER -