TY - RPRT TI - Policy Iterations for Reinforcement Learning Problems in Continuous Time and Space -- Fundamental Theory and Methods AU - Jaeyoung Lee AU - Richard S. Sutton PY - 2020 DO - 10.1016/j.automatica.2020.109421 UR - https://arxiv.org/abs/1705.03520 ID - 1705.03520 ER -