TY - RPRT TI - Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDP AU - Kefan Dong AU - Yuanhao Wang AU - Xiaoyu Chen AU - Liwei Wang PY - 2019 UR - https://arxiv.org/abs/1901.09311 ID - 1901.09311 ER -