arXiv · 2609.22234
Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models
Abstract
Instruction hierarchy (IH) alignment teaches language models to prioritize higher-level instructions when inputs conflict. While studied primarily in text-only settings, vision-language models (VLMs) introduce new challenges for IH: instructions may be embedded in images, split across modalities, visually transformed, or encountered during agentic tasks. Positing multimodal IH alignment as a reasoning problem, we train VLMs using reinforcement learning with rule-based rewards, comparing text-only, image-only, and mixed-modality supervision. We find that text-only IH training partially transfers to multimodal attacks, failing when models must decode, reconstruct, or reason over instructions across modalities. Image-based training improves robustness beyond text-only supervision, while mixed-modality training performs best overall. Importantly, the benefits generalize beyond the synthetic typographic training setting to real-image and web-agent safety tasks, while largely preserving general multimodal capability, showing that lightweight, verifiable supervision can meaningfully improve VLM robustness under adversarial, cross-modal, and interactive instruction conflicts.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nicholas Sansoterra, Zishuo Zheng, Sachin Kumar. 2026-09-03. Seeing Through Conflicts: Improving Instruction Hierarchy Alignment in Vision-Language Models. https://arxiv.org/abs/2609.22234
Cite the original work for its findings. Save a collection to share your selection of sources.