arXiv · 2608.20748
Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer
Abstract
The Visual Geometry Grounded Transformer (VGGT) enables unified feed-forward 3D reconstruction from multi-view images. However, deploying such a high-performance model may expose critical security vulnerabilities. Traditional adversarial perturbations require costly per-scene optimization, while Universal Adversarial Perturbations (UAPs) rely on a single static pattern and fail to effectively attack VGGT. To address these limitations, we propose \textbf{MVAP-G}, a multi-view adversarial perturbation generator that produces imperceptible consistent perturbations across multiple views in a single feed-forward pass. To ensure perturbation consistency across diverse scenes, we design a cross-view adversarial alignment mechanism to process multi-view images. Experiments demonstrate that MVAP-G significantly degrades VGGT performance without iterative optimization during inference. This work pioneers multi-view adversarial attacks on 3D foundation models, uncovering severe vulnerabilities and underscoring the urgent need for robust 3D vision systems. The code is available at https://github.com/qsong2001/mvap-g.
Explore related subjects
Keep this discovery
Qi Song, Ziyuan Luo, Haoliang Han, Renjie Wan. 2026-08-21. Generating Multi-view Adversarial Examples for Visual Geometry Grounded Transformer. https://arxiv.org/abs/2608.20748
Cite the original work for its findings. Save a collection to share your selection of sources.