TY - RPRT TI - Preference-based Reinforcement Learning with Finite-Time Guarantees AU - Yichong Xu AU - Ruosong Wang AU - Lin F. Yang AU - Aarti Singh AU - Artur Dubrawski PY - 2020 UR - https://arxiv.org/abs/2006.08910 ID - 2006.08910 ER -