TY - RPRT TI - Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning AU - Adrià López Escoriza AU - Nicklas Hansen AU - Stone Tao AU - Tongzhou Mu AU - Hao Su PY - 2025 UR - https://arxiv.org/abs/2503.01837 ID - 2503.01837 ER -