TY - RPRT TI - Off-Policy Reinforcement Learning with High Dimensional Reward AU - Dong Neuck Lee AU - Michael R. Kosorok PY - 2024 UR - https://arxiv.org/abs/2408.07660 ID - 2408.07660 ER -