arXiv · 2501.03142
Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies
Abstract
Deep reinforcement learning (RL) policies can demonstrate unsafe behaviors and are challenging to interpret. To address these challenges, we combine RL policy model checking--a technique for determining whether RL policies exhibit unsafe behaviors--with co-activation graph analysis--a method that maps neural network inner workings by analyzing neuron activation patterns--to gain insight into the safe RL policy's sequential decision-making. This combination lets us interpret the RL policy's inner workings for safe decision-making. We demonstrate its applicability in various experiments.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dennis Gross, Helge Spieker. 2025-01-06. Co-Activation Graph Analysis of Safety-Verified and Explainable Deep Reinforcement Learning Policies. https://arxiv.org/abs/2501.03142
Cite the original work for its findings. Save a collection to share your selection of sources.