ACL 2025finding0 citations

LCHAIM - Investigating Long Context Reasoning in Hebrew

Ehud Malul, Oriel Perets, Ziv Mor, Yigal Kassel, Elior Sulem

Abstract

Natural Language Inference (NLI) has gained significant attention recently due to its importance in understanding how machines comprehend and reason about language. While English has received tremendous interest, Morphologically Rich Languages (MRLs) like Hebrew, require more research. In this paper, we address the evaluation of Hebrew NLI models by introducing LCHAIM, a dataset designed to evaluate these models on tasks involving long premises and complex reasoning. The dataset, created by translating and validating the English ConTRoL dataset, consists of 8,325 context-hypothesis pairs that require coreferential, temporal, logical and analytical reasoning. Our experiments show the difficulty of contextual reasoning in Hebrew, as evidenced by the performance of different models. Fine-tuning the LongHero model on both the shorter premise Hebrew NLI and the LCHAIM datasets yielded a mean accuracy of 52%, that is 35% less than human performance. Similarly, Large language Models (LLMs) like Gemma-9B, Dicta-LM-2.0-7B, and GPT-4o achieved a top mean accuracy of 60.12% in few-shot setting.

BibTeX
@inproceedings{malul-etal-2025-lchaim,
    title = "{LCHAIM} - Investigating Long Context Reasoning in {H}ebrew",
    author = "Malul, Ehud  and
      Perets, Oriel  and
      Mor, Ziv  and
      Kassel, Yigal  and
      Sulem, Elior",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.413/",
    doi = "10.18653/v1/2025.findings-acl.413",
    pages = "7928--7939",
    ISBN = "979-8-89176-256-5"
}