TY - RPRT TI - Off-Policy Temporal Difference Learning for Perturbed Markov Decision Processes: Theoretical Insights and Extensive Simulations AU - Ali Forootani AU - Raffaele Iervolino AU - Massimo Tipaldi AU - Mohammad Khosravi PY - 2025 UR - https://arxiv.org/abs/2502.18415 ID - 2502.18415 ER -