TY - RPRT TI - Reinforcement Learning from Human Feedback with High-Confidence Safety Constraints AU - Yaswanth Chittepu AU - Blossom Metevier AU - Will Schwarzer AU - Austin Hoag AU - Scott Niekum AU - Philip S. Thomas PY - 2025 UR - https://arxiv.org/abs/2506.08266 ID - 2506.08266 ER -