ACL 2025finding0 citations

Financial Language Model Evaluation (FLaME)

Glenn Matlin, Mika Okamoto, Huzaifa Pardawala, Yang Yang, Sudheer Chava

Abstract

Language Models (LMs) have demonstrated impressive capabilities with core Natural Language Processing (NLP) tasks. The effectiveness of LMs for highly specialized knowledge-intensive tasks in finance remains difficult to assess due to major gaps in the methodologies of existing evaluation frameworks, which have caused an erroneous belief in a far lower bound of LMs’ performance on common Finance NLP (FinNLP) tasks. To demonstrate the potential of LMs for these FinNLP tasks, we present the first holistic benchmarking suite for Financial Language Model Evaluation (FLaME). We are the first research paper to comprehensively study LMs against ‘reasoning-reinforced’ LMs, with an empirical study of 23 foundation LMs over 20 core NLP tasks in finance. We open-source our framework software along with all data and results.

BibTeX
@inproceedings{matlin-etal-2025-financial,
    title = "Financial Language Model Evaluation ({FL}a{ME})",
    author = "Matlin, Glenn  and
      Okamoto, Mika  and
      Pardawala, Huzaifa  and
      Yang, Yang  and
      Chava, Sudheer",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.1164/",
    doi = "10.18653/v1/2025.findings-acl.1164",
    pages = "22633--22679",
    ISBN = "979-8-89176-256-5"
}
Financial Language Model Evaluation (FLaME) · ACL 2025