arXiv · 2409.10188
Enhancing RL Safety with Counterfactual LLM Reasoning
Abstract
Reinforcement learning (RL) policies may exhibit unsafe behavior and are hard to explain. We use counterfactual large language model reasoning to enhance RL policy safety post-training. We show that our approach improves and helps to explain the RL policy safety.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dennis Gross, Helge Spieker. 2024-09-16. Enhancing RL Safety with Counterfactual LLM Reasoning. https://arxiv.org/abs/2409.10188
Cite the original work for its findings. Save a collection to share your selection of sources.