TY - RPRT TI - Smoothed functional-based gradient algorithms for off-policy reinforcement learning: A non-asymptotic viewpoint AU - Nithia Vijayan AU - Prashanth L. A PY - 2024 UR - https://arxiv.org/abs/2101.02137 ID - 2101.02137 ER -