TY - RPRT TI - Bilevel reinforcement learning via the development of hyper-gradient without lower-level convexity AU - Yan Yang AU - Bin Gao AU - Ya-xiang Yuan PY - 2025 UR - https://arxiv.org/abs/2405.19697 ID - 2405.19697 ER -