SearcharxivSearch

arXiv subjects

Ligong Zhang

Publications and source records attributed to Ligong Zhang.

2 recordsLinked to original sources

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. Built on DriveLM-style multi-view visual question answering, the challenge represents each scene with six synchronized camera images and a structured collection of driving-related question-answer pairs. Participants generate adversarial images and suffix-only textual perturbations that induce model responses to deviate from reference answers while preserving image fidelity and limiting textual cost. The competition comprises two phases, with Phase II adding a hidden black-box model to assess transferability. We describe the task design, submission rules, evaluation protocol, and leaderboard results, and then examine five leading submissions for which technical reports were available. Across these reports, several recurring patterns emerge: image-side attacks are favored by the suffix penalty; scene-level, multi-view optimization is more effective than treating views in isolation; QA types and graph structure provide useful priors for allocating attack budget; feature-space objectives can improve black-box transfer; and typographic content embedded in camera images exposes a persistent vulnerability in driving VLAs. These findings provide a practical reference for future robustness evaluation and defense design in multimodal autonomous-driving systems.

cs.CV

Deep Learning Accelerated First-Principles Quantum Transport Simulations at Nonequilibrium State

The non-equilibrium Green's function method combined with density functional theory (NEGF-DFT) provides a rigorous framework for simulating nanoscale electronic transport, but its computational cost scales steeply with system size. Recent artificial intelligence (AI) approaches have sought to accelerate such simulations, yet most rely on conventional machine learning, lack atomic resolution, struggle to extrapolate to larger systems, and cannot predict multiple properties simultaneously. Here we introduce DeepQT, a deep-learning framework that integrates graph neural networks with transformer architectures to enable multi-property predictions of electronic structure and transport without manual feature engineering. By learning key intermediate quantities of NEGF-DFT, the equilibrium Hamiltonian and the non-equilibrium total potential difference, DeepQT reconstructs Hamiltonians under both equilibrium and bias conditions, yielding accurate transport predictions. Leveraging the principle of electronic nearsightedness, DeepQT generalizes from small training systems to much larger ones with high fidelity. Benchmarks on graphene, MoS2, and silicon diodes with varied defects and dopants show that DeepQT achieves first-principles accuracy while reducing computational cost by orders of magnitude. This scalable, transferable framework advances AI-assisted quantum transport, offering a powerful tool for next-generation nanoelectronic device design.

cond-mat.mes-hall