TY - RPRT TI - Total stochastic gradient algorithms and applications in reinforcement learning AU - Paavo Parmas PY - 2019 UR - https://arxiv.org/abs/1902.01722 ID - 1902.01722 ER -