TY - RPRT TI - Learning Self-Correctable Policies and Value Functions from Demonstrations with Negative Sampling AU - Yuping Luo AU - Huazhe Xu AU - Tengyu Ma PY - 2019 UR - https://arxiv.org/abs/1907.05634 ID - 1907.05634 ER -