TY - RPRT TI - Combining policy gradient and Q-learning AU - Brendan O'Donoghue AU - Remi Munos AU - Koray Kavukcuoglu AU - Volodymyr Mnih PY - 2017 UR - https://arxiv.org/abs/1611.01626 ID - 1611.01626 ER -