TY - RPRT TI - Trajectory-Based Off-Policy Deep Reinforcement Learning AU - Andreas Doerr AU - Michael Volpp AU - Marc Toussaint AU - Sebastian Trimpe AU - Christian Daniel PY - 2019 UR - https://arxiv.org/abs/1905.05710 ID - 1905.05710 ER -