arXiv · 2510.01038
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
Abstract
Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants. This raises doubts about the quality of the explanations. In this paper, we introduce a novel forward pass paradigm, Activation-Deactivation (AD), which obviates the need for perturbation of the input. AD replaces perturbation of input features with switching off parts of the model corresponding to to the intended perturbations. We implement ConvAD, an AD approximation algorithm for CNNs. ConvAD is a drop-in mechanism that can be easily added to any trained CNN and, without any additional training, generates more robust and more transferable explanations. We provide evaluation results across multiple architectures, datasets, methods and perturbation strategies, demonstrating the superior quality of ConvAD compared to the SOTA.
Explore related subjects
Keep this discovery
Akchunya Chanchal, David A. Kelly, Hana Chockler. 2025-10-01. Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI. https://arxiv.org/abs/2510.01038
Cite the original work for its findings. Save a collection to share your selection of sources.