arXiv · 2609.34145
Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving
Abstract
Retrieval-augmented generation (RAG) gives vision--language driving systems access to external safety knowledge, yet a retrieved risk rule may be relevant without applying to the current scene. A vision--language model (VLM) receiving such knowledge must ground objects, bind entities across time, and verify relations before deciding how to act, leaving the support for risk conclusions implicit. We address this relevance--applicability gap with a Driving-Risk Knowledge Graph (DRKG) and Semantic Web Rule Language (SWRL) reasoning stage before VLM decision-making. Structured perception instantiates scene facts, from which SWRL rules derive events and directed risk relations when their antecedents are jointly satisfied. Recognized events, bound risk relations, and semantic descriptions of activated rules form compact evidence that conditions the VLM and diffusion planner. In matched comparisons on nuReasoning, our method improved the nuReasoning planning score (NPS) by 1.30 points and the non-at-fault collision score (NC) by 2.76 points over the relevance retrieval-based baseline. These gains indicate that scene-applicable risk evidence improves safety-weighted planning relative to semantically retrieved risk knowledge.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jiaxin Liu, Ruilin Yu, Liang Peng, Jingkai Wang, Chengxiang Zhao, Zhenxin Zhu, Bing Wang, Guang Chen, Hangjun Ye, Hong Wang, Jun Li. 2026-09-28. Beyond Retrieval Relevance: Scene-Grounded Risk Entailment for Vision-Language Driving. https://arxiv.org/abs/2609.34145
Cite the original work for its findings. Save a collection to share your selection of sources.