arXiv · 2508.05600
Provable one-poison backdoor attacks on linear models and ReLU neural networks
Abstract
Backdoor poisoning attacks are a threat to machine learning models that are trained on data collected from untrusted sources; these attacks enable attackers to inject malicious behavior into the model that can be triggered by specially crafted inputs. Prior work has established bounds on the success of backdoor attacks and their impact on the benign learning task, however, an open question is what amount of poison data is needed for a successful backdoor attack. Typical attacks either use few samples but need much information about the data points, or need to poison many data points. In this paper, we show that an adversary can mount a one-poison backdoor attack without knowledge of individual training data, requiring only coarse geometric bounds of the input space and training parameters. We identify provably sufficient conditions that allow an adversary with one poison sample with high probability to inject a backdoor into linear models and MLPs. We show that our backdoor has zero backdooring error and the injection does not significantly impact the benign learning task performance.
Explore related subjects
Keep this discovery
Thorsten Peinemann, Paula Arnold, Sebastian Berndt, Thomas Eisenbarth, Esfandiar Mohammadi. 2025-08-07. Provable one-poison backdoor attacks on linear models and ReLU neural networks. https://arxiv.org/abs/2508.05600
Cite the original work for its findings. Save a collection to share your selection of sources.