arXiv · 2412.13377
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
Abstract
This paper introduces DateLogicQA, a benchmark with 190 questions covering diverse date formats, temporal contexts, and reasoning types. We propose the Semantic Integrity Metric to assess tokenization quality and analyse two biases: Representation-Level Bias, affecting embeddings, and Logical-Level Bias, influencing reasoning outputs. Our findings provide a comprehensive evaluation of LLMs' capabilities and limitations in temporal reasoning, highlighting key challenges in handling temporal data accurately.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Gagan Bhatia, MingZe Tang, Cristina Mahanta, Madiha Kazi. 2024-12-17. DateLogicQA: Benchmarking Temporal Biases in Large Language Models. https://arxiv.org/abs/2412.13377
Cite the original work for its findings. Save a collection to share your selection of sources.