TY - RPRT TI - Finite-Time Performance Bounds and Adaptive Learning Rate Selection for Two Time-Scale Reinforcement Learning AU - Harsh Gupta AU - R. Srikant AU - Lei Ying PY - 2019 UR - https://arxiv.org/abs/1907.06290 ID - 1907.06290 ER -