TY - RPRT TI - A Provable Approach for End-to-End Safe Reinforcement Learning AU - Akifumi Wachi AU - Kohei Miyaguchi AU - Takumi Tanabe AU - Rei Sato AU - Youhei Akimoto PY - 2025 UR - https://arxiv.org/abs/2505.21852 ID - 2505.21852 ER -