TY - RPRT TI - Learning from Suboptimal Demonstration via Self-Supervised Reward Regression AU - Letian Chen AU - Rohan Paleja AU - Matthew Gombolay PY - 2020 UR - https://arxiv.org/abs/2010.11723 ID - 2010.11723 ER -