TY - RPRT TI - Learning Reward Functions from Diverse Sources of Human Feedback: Optimally Integrating Demonstrations and Preferences AU - Erdem Bıyık AU - Dylan P. Losey AU - Malayandi Palan AU - Nicholas C. Landolfi AU - Gleb Shevchuk AU - Dorsa Sadigh PY - 2021 UR - https://arxiv.org/abs/2006.14091 ID - 2006.14091 ER -