TY - RPRT TI - Near-optimal Regret Using Policy Optimization in Online MDPs with Aggregate Bandit Feedback AU - Tal Lancewicki AU - Yishay Mansour PY - 2025 UR - https://arxiv.org/abs/2502.04004 ID - 2502.04004 ER -