TY - RPRT TI - Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning AU - Alex Ayoub AU - Kaiwen Wang AU - Vincent Liu AU - Samuel Robertson AU - James McInerney AU - Dawen Liang AU - Nathan Kallus AU - Csaba Szepesvári PY - 2024 UR - https://arxiv.org/abs/2403.05385 ID - 2403.05385 ER -