TY - RPRT TI - LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration AU - Ruiyu Qiu AU - Rui Wang AU - Guanghui Yang AU - Xiang Li AU - Zhijiang Shao PY - 2025 UR - https://arxiv.org/abs/2511.08339 ID - 2511.08339 ER -