TY - RPRT TI - Two Timescale Stochastic Approximation with Controlled Markov noise and Off-policy temporal difference learning AU - Prasenjit Karmakar AU - Shalabh Bhatnagar PY - 2017 UR - https://arxiv.org/abs/1503.09105 ID - 1503.09105 ER -