TY - RPRT TI - Sharp Variance-Dependent Bounds in Reinforcement Learning: Best of Both Worlds in Stochastic and Deterministic Environments AU - Runlong Zhou AU - Zihan Zhang AU - Simon S. Du PY - 2023 UR - https://arxiv.org/abs/2301.13446 ID - 2301.13446 ER -