arXiv · 2610.10423
GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks
Abstract
Adversarial example detectors are often tied to the classifier backbone they were trained on, limiting reuse when the protected model is replaced or upgraded. Directly transferring such detectors across backbones is challenging because different networks generally produce incompatible internal representations. We propose GraphRectify, a graph-based framework for transferring adversarial image detectors across classifier backbones. GraphRectify learns a structured representation of intermediate classifier features and adapts representations from a new backbone to the detector learned on the original model, enabling detector reuse. We evaluate GraphRectify across multiple datasets, backbone architectures, and adversarial attacks, including detector-aware adaptive attacks that jointly target the classifier and detector. Across the complete evaluation matrix, GraphRectify achieves higher aggregate ROC-AUC than training a detector from scratch on the new backbone and the evaluated transfer ablations. The gains are particularly strong for transfers between different backbone families and when sufficient data are available. In contrast, training from scratch remains competitive in the most data-limited settings. These results show that adversarial detection knowledge can transfer effectively across heterogeneous classifier architectures rather than being relearned whenever the protected backbone changes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Arash Vashagh, Roozbeh Razavi-Far. 2026-10-07. GraphRectify: Graph-Based Transfer of Adversarial Example Detectors Across Neural Networks. https://arxiv.org/abs/2610.10423
Cite the original work for its findings. Save a collection to share your selection of sources.