TY - RPRT TI - Batch Active Learning of Reward Functions from Human Preferences AU - Erdem Bıyık AU - Nima Anari AU - Dorsa Sadigh PY - 2024 UR - https://arxiv.org/abs/2402.15757 ID - 2402.15757 ER -