TY - RPRT TI - The Local Optimality of Reinforcement Learning by Value Gradients, and its Relationship to Policy Gradient Learning AU - Michael Fairbank AU - Eduardo Alonso PY - 2011 UR - https://arxiv.org/abs/1101.0428 ID - 1101.0428 ER -