arXiv · 2601.03605
AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification
Abstract
Despite the significant advancements of Large Language Models (LLMs), their factuality remains a critical challenge, creating a growing need for more nuanced factuality verification. Existing factuality verification methods do not capture graded judgments, even though factuality is better understood as a spectrum rather than a binary of right and wrong. To bridge this gap, we focus on graded factuality verification and propose AEScorer, an agentic evidence-grounded framework with two stages: agentic evidence acquisition and graded scoring. AEScorer first gathers and refines external evidence through agentic search, and then predicts a scalar factuality score to distinguish nuanced differences in factual correctness. We further construct GradedVeriBench, a benchmark for graded factuality verification spanning both general and multi-hop question answering. Experimental results on GradedVeriBench show that AEScorer substantially outperforms existing methods across both settings, demonstrating the value of coupling targeted evidence acquisition with graded scoring.
Explore related subjects
Keep this discovery
Hui Huang, Muyun Yang, Yuki Arase. 2026-01-07. AEScorer: An Agentic Evidence-Grounded Framework for Graded Factuality Verification. https://arxiv.org/abs/2601.03605
Cite the original work for its findings. Save a collection to share your selection of sources.