TY - RPRT TI - Online Reinforcement Learning for Real-Time Exploration in Continuous State and Action Markov Decision Processes AU - Ludovic Hofer AU - Hugo Gimbert PY - 2016 UR - https://arxiv.org/abs/1612.03780 ID - 1612.03780 ER -