TY - RPRT TI - Statistically Efficient Variance Reduction with Double Policy Estimation for Off-Policy Evaluation in Sequence-Modeled Reinforcement Learning AU - Hanhan Zhou AU - Tian Lan AU - Vaneet Aggarwal PY - 2023 UR - https://arxiv.org/abs/2308.14897 ID - 2308.14897 ER -