TY - RPRT TI - Multiple-Step Greedy Policies in Online and Approximate Reinforcement Learning AU - Yonathan Efroni AU - Gal Dalal AU - Bruno Scherrer AU - Shie Mannor PY - 2018 UR - https://arxiv.org/abs/1805.07956 ID - 1805.07956 ER -