TY - RPRT TI - Anchor-Changing Regularized Natural Policy Gradient for Multi-Objective Reinforcement Learning AU - Ruida Zhou AU - Tao Liu AU - Dileep Kalathil AU - P. R. Kumar AU - Chao Tian PY - 2022 UR - https://arxiv.org/abs/2206.05357 ID - 2206.05357 ER -