TY - RPRT TI - Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial Coverage AU - Jose Blanchet AU - Miao Lu AU - Tong Zhang AU - Han Zhong PY - 2023 UR - https://arxiv.org/abs/2305.09659 ID - 2305.09659 ER -