arXiv · 2503.07144
MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark
Abstract
Machine Reading Comprehension (MRC) is an essential task in evaluating natural language understanding. Existing MRC datasets primarily assess specific aspects of reading comprehension (RC), lacking a comprehensive MRC benchmark. To fill this gap, we first introduce a novel taxonomy that categorizes the key capabilities required for RC. Based on this taxonomy, we construct MRCEval, an MRC benchmark that leverages advanced Large Language Models (LLMs) as both sample generators and selection judges. MRCEval is a comprehensive, challenging and accessible benchmark designed to assess the RC capabilities of LLMs thoroughly, covering 13 distinct RC skills with a total of 2.1K high-quality multi-choice questions. We perform an extensive evaluation of 28 widely used open-source and proprietary models, highlighting that MRC continues to present significant challenges even in the era of LLMs.
Explore related subjects
Keep this discovery
Shengkun Ma, Hao Peng, Lei Hou, Juanzi Li. 2025-03-10. MRCEval: A Comprehensive, Challenging and Accessible Machine Reading Comprehension Benchmark. https://arxiv.org/abs/2503.07144
Cite the original work for its findings. Save a collection to share your selection of sources.