arXiv · 2602.04711
Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention
Abstract
Retrieval Augmented Generation (RAG) is a highly effective paradigm for keeping LLM-based responses up-to-date and reducing the likelihood of hallucinations. Yet, RAG was recently shown to be quite vulnerable to corpus knowledge poisoning: an attacker injects misleading documents to the corpus to steer an LLM's output to an undesired response. We argue that the standard causal attention mechanism in LLMs enables harmful cross-document interactions, specifically in cases of attacks. Accordingly, we introduce a novel defense approach for RAG: Sparse Document Attention RAG (SDAG). This is a block-sparse attention mechanism that disallows cross-attention between retrieved documents. SDAG requires a minimal inference-time change to the attention mask. We present an empirical evaluation of LLM-based question answering (QA) with a variety of attack strategies on RAG. We show that our SDAG method substantially outperforms the standard causal attention mechanism. We further demonstrate the clear merits of integrating SDAG with state-of-the-art RAG defense methods. Specifically, the integration results in performance that is statistically significantly better than the state-of-the-art.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sagie Dekel, Moshe Tennenholtz, Oren Kurland. 2026-02-04. Addressing Corpus Knowledge Poisoning Attacks on RAG Using Sparse Attention. https://arxiv.org/abs/2602.04711
Cite the original work for its findings. Save a collection to share your selection of sources.