TY - RPRT TI - The Divergence of Reinforcement Learning Algorithms with Value-Iteration and Function Approximation AU - Michael Fairbank AU - Eduardo Alonso PY - 2012 UR - https://arxiv.org/abs/1107.4606 ID - 1107.4606 ER -