arXiv · 2003.05748
Explaining Away Attacks Against Neural Networks
Abstract
We investigate the problem of identifying adversarial attacks on image-based neural networks. We present intriguing experimental results showing significant discrepancies between the explanations generated for the predictions of a model on clean and adversarial data. Utilizing this intuition, we propose a framework which can identify whether a given input is adversarial based on the explanations given by the model. Code for our experiments can be found here: https://github.com/seansaito/Explaining-Away-Attacks-Against-Neural-Networks.
Explore related subjects
Keep this discovery
Sean Saito, Jin Wang. 2020-03-06. Explaining Away Attacks Against Neural Networks. https://arxiv.org/abs/2003.05748
Cite the original work for its findings. Save a collection to share your selection of sources.