arXiv · 2601.12304
A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models
Abstract
Vision-language pre-training (VLP) models are vulnerable to adversarial examples, particularly in black-box scenarios. Existing multimodal attacks often suffer from limited perturbation diversity and unstable multi-stage pipelines. To address these challenges, we propose 2S-GDA, a two-stage globally-diverse attack framework. The proposed method first introduces textual perturbations through a globally-diverse strategy by combining candidate text expansion with globally-aware replacement. To enhance visual diversity, image-level perturbations are generated using multi-scale resizing and block-shuffle rotation. Extensive experiments on VLP models demonstrate that 2S-GDA consistently improves attack success rates over state-of-the-art methods, with gains of up to 11.17\% in black-box settings. Our framework is modular and can be easily combined with existing methods to further enhance adversarial transferability.
Explore related subjects
Keep this discovery
Wutao Chen, Huaqin Zou, Chen Wan, Lifeng Huang. 2026-01-18. A Two-Stage Globally-Diverse Adversarial Attack for Vision-Language Pre-training Models. https://arxiv.org/abs/2601.12304
Cite the original work for its findings. Save a collection to share your selection of sources.