TY - RPRT TI - The Mirage of Action-Dependent Baselines in Reinforcement Learning AU - George Tucker AU - Surya Bhupatiraju AU - Shixiang Gu AU - Richard E. Turner AU - Zoubin Ghahramani AU - Sergey Levine PY - 2018 UR - https://arxiv.org/abs/1802.10031 ID - 1802.10031 ER -