TY - RPRT TI - Learning Adversarial Low-rank Markov Decision Processes with Unknown Transition and Full-information Feedback AU - Canzhe Zhao AU - Ruofeng Yang AU - Baoxiang Wang AU - Xuezhou Zhang AU - Shuai Li PY - 2023 UR - https://arxiv.org/abs/2311.07876 ID - 2311.07876 ER -