arXiv · 2609.23596
StyleAT: Defending Face Recognition Against Semantic Attacks
Abstract
With face-recognition models now embedded in everyday authentication and surveillance, recent works have pinpointed a critical weakness: these models remain acutely vulnerable to adversarial semantic edits. I.e., adversarially produced semantic alterations to the input, such as slight aging or pose changes, can induce misclassifications. Certain existing attacks are powerful, but they can be computationally costly, rendering them inadequate for developing defenses (e.g., through adversarial training). To fill the gap, we introduce BoundStyle, a potent semantic attack operating in StyleGAN's rich latent space to maximize misclassification rates. Notably, BoundStyle achieves high attack success rates while being ${\sim}{\times}9.5$ faster than existing state-of-the-art attacks, making it suitable for adversarial training. Building on BoundStyle, we develop StyleAT, an efficient adversarial training scheme that incorporates low-budget attack variants yet defends against stronger and unseen semantic attacks. We evaluate on two datasets unseen during training and seven models, and find that StyleAT boosts robust accuracy against state-of-the-art attacks and outperforms common defenses in various settings.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ben Shapira, Roi Cohen, Shang-Tse Chen, Mahmood Sharif. 2026-09-20. StyleAT: Defending Face Recognition Against Semantic Attacks. https://arxiv.org/abs/2609.23596
Cite the original work for its findings. Save a collection to share your selection of sources.