TY - RPRT TI - Reward Estimation for Variance Reduction in Deep Reinforcement Learning AU - Joshua Romoff AU - Peter Henderson AU - Alexandre Piché AU - Vincent Francois-Lavet AU - Joelle Pineau PY - 2018 UR - https://arxiv.org/abs/1805.03359 ID - 1805.03359 ER -