TY - RPRT TI - Sample-Efficient Preference-based Reinforcement Learning with Dynamics Aware Rewards AU - Katherine Metcalf AU - Miguel Sarabia AU - Natalie Mackraz AU - Barry-John Theobald PY - 2024 UR - https://arxiv.org/abs/2402.17975 ID - 2402.17975 ER -