TY - RPRT TI - Hindsight Preference Replay Improves Preference-Conditioned Multi-Objective Reinforcement Learning AU - Jonaid Shianifar AU - Michael Schukat AU - Karl Mason PY - 2026 UR - https://arxiv.org/abs/2601.11604 ID - 2601.11604 ER -