arXiv · 2511.16134
Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents
Abstract
Table Extraction (TE) consists in extracting tables from PDF documents, in a structured format enabling automatic processing. While numerous TE tools exist, the variety of methods and techniques makes it difficult for users to choose the most appropriate one. We propose a novel benchmark for assessing end-to-end TE methods (from PDF to the final table) over 86k pages. We contribute an analysis of TE evaluation metrics, and a novel, rigorous evaluation process, which allows scoring each TE sub-task as well as end-to-end TE, and captures model uncertainty. Along with prior datasets, our benchmark comprises two new heterogeneous datasets of 39k samples. We run our benchmark on diverse models, including off-the-shelf libraries, tools, computer vision-based models and modern approaches using general and specialized vision language models. The results demonstrate that TE remains challenging: current methods suffer from a lack of generalizability when facing heterogeneous data, and from limitations in robustness and interpretability.
Explore related subjects
Keep this discovery
Marijan Soric, Cécile Gracianne, Ioana Manolescu, Pierre Senellart. 2025-11-20. Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents. https://doi.org/10.1145/3770855.3817462
Cite the original work for its findings. Save a collection to share your selection of sources.