TY - RPRT TI - Near-Optimal Reinforcement Learning with Self-Play under Adaptivity Constraints AU - Dan Qiao AU - Yu-Xiang Wang PY - 2024 UR - https://arxiv.org/abs/2402.01111 ID - 2402.01111 ER -