arXiv · 2605.10862
RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems
Abstract
This paper demonstrates RUBEN, an interactive tool for discovering minimal rules to explain the outputs of retrieval-augmented large language models (LLMs) in data-driven applications. We leverage novel pruning strategies to efficiently identify a minimal set of rules that subsume all others. We further demonstrate novel applications of these rules for LLM safety, specifically to test the resiliency of safety training and effectiveness of adversarial prompt injections.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Joel Rorseth, Parke Godfrey, Lukasz Golab, Divesh Srivastava, Jarek Szlichta. 2026-05-11. RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems. https://arxiv.org/abs/2605.10862
Cite the original work for its findings. Save a collection to share your selection of sources.