arXiv · 1809.02397
Detecting Potential Local Adversarial Examples for Human-Interpretable Defense
Abstract
Machine learning models are increasingly used in the industry to make decisions such as credit insurance approval. Some people may be tempted to manipulate specific variables, such as the age or the salary, in order to get better chances of approval. In this ongoing work, we propose to discuss, with a first proposition, the issue of detecting a potential local adversarial example on classical tabular data by providing to a human expert the locally critical features for the classifier's decision, in order to control the provided information and avoid a fraud.
Explore related subjects
Keep this discovery
Xavier Renard, Thibault Laugel, Marie-Jeanne Lesot, Christophe Marsala, Marcin Detyniecki. 2018-09-07. Detecting Potential Local Adversarial Examples for Human-Interpretable Defense. https://arxiv.org/abs/1809.02397
Cite the original work for its findings. Save a collection to share your selection of sources.