TY - RPRT TI - Approximating two value functions instead of one: towards characterizing a new family of Deep Reinforcement Learning algorithms AU - Matthia Sabatelli AU - Gilles Louppe AU - Pierre Geurts AU - Marco A. Wiering PY - 2019 UR - https://arxiv.org/abs/1909.01779 ID - 1909.01779 ER -