arXiv · 2504.16188
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking
Abstract
We introduce FinNLI, a benchmark dataset for Financial Natural Language Inference (FinNLI) across diverse financial texts like SEC Filings, Annual Reports, and Earnings Call transcripts. Our dataset framework ensures diverse premise-hypothesis pairs while minimizing spurious correlations. FinNLI comprises 21,304 pairs, including a high-quality test set of 3,304 instances annotated by finance experts. Evaluations show that domain shift significantly degrades general-domain NLI performance. The highest Macro F1 scores for pre-trained (PLMs) and large language models (LLMs) baselines are 74.57% and 78.62%, respectively, highlighting the dataset's difficulty. Surprisingly, instruction-tuned financial LLMs perform poorly, suggesting limited generalizability. FinNLI exposes weaknesses in current LLMs for financial reasoning, indicating room for improvement.
Explore related subjects
Keep this discovery
Jabez Magomere, Elena Kochkina, Samuel Mensah, Simerjot Kaur, Charese H. Smiley. 2025-04-22. FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking. https://arxiv.org/abs/2504.16188
Cite the original work for its findings. Save a collection to share your selection of sources.