arXiv · 2508.11758
Can we Evaluate RAGs with Synthetic Data?
Abstract
We investigate whether synthetic question-answer (QA) data generated by large language models (LLMs) can serve as an effective proxy for human-labeled benchmarks when the latter is unavailable. We assess the reliability of synthetic benchmarks across two experiments: one varying retriever parameters while keeping the generator fixed, and another varying the generator with fixed retriever parameters. Across four datasets, of which two open-domain and two proprietary, we find that synthetic benchmarks reliably rank the RAGs varying in terms of retriever configuration, aligning well with human-labeled benchmark baselines. However, they do not consistently produce reliable RAG rankings when comparing generator architectures. The breakdown possibly arises from a combination of task mismatch between the synthetic and human benchmarks, and stylistic bias favoring certain generators.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jonas van Elburg, Peter van der Putten, Maarten Marx. 2025-08-15. Can we Evaluate RAGs with Synthetic Data?. https://arxiv.org/abs/2508.11758
Cite the original work for its findings. Save a collection to share your selection of sources.