arXiv · 2606.25721
Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution
Abstract
Retrieval-Augmented Generation (RAG) systems are vulnerable to corpus poisoning attacks that manipulate model outputs through malicious retrieved documents. Existing detection methods typically rely on auxiliary classifiers or additional LLM-based verification, introducing substantial computational overhead. We present TRACE, a lightweight detection framework that identifies poisoning attacks by tracing answer-related tokens through token influence attribution. TRACE first discovers recurrent high-influence keywords across retrieved documents and then performs a secondary verification to confirm their influence on model predictions. Experiments on three QA benchmarks and six LLMs demonstrate strong detection performance while simultaneously uncovering attacker-specified target answers.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yan-Lun Chen, Pin-Yu Chen, Chia-Mu Yu, Ying-Dar Lin, Yu-Sung Wu, Wei-Bin Lee. 2026-06-24. Tracing Target Answers in Poisoned Retrieval Corpora via Token Influence Attribution. https://arxiv.org/abs/2606.25721
Cite the original work for its findings. Save a collection to share your selection of sources.