arXiv · 2505.13538
RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines
Abstract
Retrieval-Augmented Generation (RAG) systems couple large language models with external knowledge, yet most evaluation methods report aggregate scores that reveal whether a pipeline underperforms but not where or why. We introduce RAGXplain, an evaluation framework that translates performance metrics into actionable guidance. RAGXplain structures evaluation around a 'Metric Diamond' connecting user input, retrieved context, generated answer, and (when available) ground truth via six diagnostic dimensions. It uses LLM reasoning to produce natural-language failure-mode explanations and prioritized interventions. Across five QA benchmarks, applying RAGXplain's recommendations in a single human-guided pass consistently improves RAG pipeline performance across multiple metrics. We release RAGXplain as open source to support reproducibility and community adoption.
Explore related subjects
Keep this discovery
Dvir Cohen, Tamir Houri, Lin Burg, Gilad Barkan. 2025-05-18. RAGXplain: From Explainable Evaluation to Actionable Guidance of RAG Pipelines. https://arxiv.org/abs/2505.13538
Cite the original work for its findings. Save a collection to share your selection of sources.