arXiv · 1908.01469
Automated Detection System for Adversarial Examples with High-Frequency Noises Sieve
Abstract
Deep neural networks are being applied in many tasks with encouraging results, and have often reached human-level performance. However, deep neural networks are vulnerable to well-designed input samples called adversarial examples. In particular, neural networks tend to misclassify adversarial examples that are imperceptible to humans. This paper introduces a new detection system that automatically detects adversarial examples on deep neural networks. Our proposed system can mostly distinguish adversarial samples and benign images in an end-to-end manner without human intervention. We exploit the important role of the frequency domain in adversarial samples and propose a method that detects malicious samples in observations. When evaluated on two standard benchmark datasets (MNIST and ImageNet), our method achieved an out-detection rate of 99.7 - 100% in many settings.
Explore related subjects
Keep this discovery
Dang Duy Thang, Toshihiro Matsui. 2019-08-05. Automated Detection System for Adversarial Examples with High-Frequency Noises Sieve. https://arxiv.org/abs/1908.01469
Cite the original work for its findings. Save a collection to share your selection of sources.