arXiv · 2506.18421
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Abstract
The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap by presenting a comprehensive table reasoning benchmark, TReB. Firstly, we propose a taxonomy to systematically measure both shallow table understanding abilities and deep table reasoning abilities, covering a total of 26 sub-tasks. We then construct a high quality dataset through a dedicated data processing and synthesis procedure. Based on these well-constructed samples, we design an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes. Experimental results with our data and framework reveal that existing LLMs still have significant room for improvement in addressing the complex and real world table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on https://huggingface.co/datasets/JT-LM/JIUTIAN-TReB, and the framework on https://github.com/JT-LM/jiutian-treb.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ce Li, Xiaofan Liu, Zhiyan Song, Ce Chi, Boshen Shi, Chen Zhao, Guanguang Chang, Zhendong Wang, Kexin Yang, Xing Wang, Chao Deng, Junlan Feng. 2025-06-23. TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models. https://doi.org/10.1145/3805712.3808617
Cite the original work for its findings. Save a collection to share your selection of sources.