TY - RPRT TI - When is Offline Policy Selection Sample Efficient for Reinforcement Learning? AU - Vincent Liu AU - Prabhat Nagarajan AU - Andrew Patterson AU - Martha White PY - 2026 UR - https://arxiv.org/abs/2312.02355 ID - 2312.02355 ER -