NAACL 2025findings1 citations

FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking

Jabez Magomere, Elena Kochkina, Samuel Mensah, Simerjot Kaur, Charese Smiley

Abstract

We introduce FinNLI, a benchmark dataset for Financial Natural Language Inference (FinNLI) across diverse financial texts like SEC Filings, Annual Reports, and Earnings Call transcripts. Our dataset framework ensures diverse premise-hypothesis pairs while minimizing spurious correlations. FinNLI comprises 21,304 pairs, including a high-quality test set of 3,304 instances annotated by finance experts. Evaluations show that domain shift significantly degrades general-domain NLI performance. The highest Macro F1 scores for pre-trained (PLMs) and large language models (LLMs) baselines are 74.57% and 78.62%, respectively, highlighting the dataset’s difficulty. Surprisingly, instruction-tuned financial LLMs perform poorly, suggesting limited generalizability. FinNLI exposes weaknesses in current LLMs for financial reasoning, indicating room for improvement.

BibTeX
@inproceedings{magomere-etal-2025-finnli,
    title = "{F}in{NLI}: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking",
    author = "Magomere, Jabez  and
      Kochkina, Elena  and
      Mensah, Samuel  and
      Kaur, Simerjot  and
      Smiley, Charese",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Findings of the Association for Computational Linguistics: NAACL 2025",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-naacl.257/",
    pages = "4545--4568",
    ISBN = "979-8-89176-195-7"
}
FinNLI: Novel Dataset for Multi-Genre Financial Natural Language Inference Benchmarking · NAACL 2025