arXiv · 2506.14397
Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding
Abstract
Negation is a fundamental linguistic phenomenon that poses ongoing challenges for Large Language Models (LLMs), particularly in tasks requiring deep semantic understanding. Current benchmarks often treat negation as a minor detail within broader tasks, such as natural language inference. Consequently, there is a lack of benchmarks specifically designed to evaluate comprehension of negation. In this work, we introduce Thunder-NUBench, a novel benchmark explicitly created to assess sentence-level understanding of negation in LLMs. Thunder-NUBench goes beyond merely identifying surface-level cues by contrasting standard negation with structurally diverse alternatives, such as local negation, contradiction, and paraphrase. This benchmark includes manually curated sentence-negation pairs and a multiple-choice dataset, allowing for a comprehensive evaluation of models' understanding of negation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yeonkyoung So, Gyuseong Lee, Sungmok Jung, Joonhak Lee, JiA Kang, Sangho Kim, Jaejin Lee. 2025-06-17. Thunder-NUBench: A Benchmark for LLMs' Sentence-Level Negation Understanding. https://doi.org/10.18653/v1%2F2026.findings-eacl.250
Cite the original work for its findings. Save a collection to share your selection of sources.