TY - RPRT TI - Deep Q-learning from Demonstrations AU - Todd Hester AU - Matej Vecerik AU - Olivier Pietquin AU - Marc Lanctot AU - Tom Schaul AU - Bilal Piot AU - Dan Horgan AU - John Quan AU - Andrew Sendonaris AU - Gabriel Dulac-Arnold AU - Ian Osband AU - John Agapiou AU - Joel Z. Leibo AU - Audrunas Gruslys PY - 2017 UR - https://arxiv.org/abs/1704.03732 ID - 1704.03732 ER -