TY - RPRT TI - Meta-Reinforcement Learning with Universal Policy Adaptation: Provable Near-Optimality under All-task Optimum Comparator AU - Siyuan Xu AU - Minghui Zhu PY - 2024 UR - https://arxiv.org/abs/2410.09728 ID - 2410.09728 ER -