arXiv · 2001.05207
A Formal Approach to Explainability
Abstract
We regard explanations as a blending of the input sample and the model's output and offer a few definitions that capture various desired properties of the function that generates these explanations. We study the links between these properties and between explanation-generating functions and intermediate representations of learned models and are able to show, for example, that if the activations of a given layer are consistent with an explanation, then so do all other subsequent layers. In addition, we study the intersection and union of explanations as a way to construct new explanations.
Explore related subjects
Keep this discovery
Lior Wolf, Tomer Galanti, Tamir Hazan. 2020-01-15. A Formal Approach to Explainability. https://arxiv.org/abs/2001.05207
Cite the original work for its findings. Save a collection to share your selection of sources.