arXiv · 2510.17853
CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation
Abstract
Large Language Models (LLMs) have emerged as powerful assistants for scientific writing. However, concerns remain about the quality and reliability of the generated text, including citation accuracy and faithfulness. While most recent work relies on methods such as LLM-as-a-Judge, the reliability of LLM-as-a-Judge alone is also in doubt. In this work, we reframe citation evaluation as a problem of citation attribution alignment, which assesses whether LLM-generated citations match those a human author would include for the same text. We propose CiteGuard, a retrieval-aware agent framework designed to provide more faithful grounding for citation validation. CiteGuard improves over the prior baseline by 10 percentage points and achieves up to 68.1% accuracy on the CiteME benchmark, approaching human performance (69.2%). It also identifies alternative valid citations and demonstrates generalization ability for cross-domain citation attribution. Our code is available at https://github.com/KathCYM/CiteGuard.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yee Man Choi, Xuehang Guo, Yi R. Fung, Qingyun Wang. 2025-10-15. CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation. https://arxiv.org/abs/2510.17853
Cite the original work for its findings. Save a collection to share your selection of sources.