EMNLP 20250 citations

Assessing the Sensitivity and Alignment of FOL Closeness Metrics

Ramya Keerthy Thatikonda, Wray Buntine, Ehsan Shareghi

Abstract

The recent successful paradigm of solving logical reasoning problems with tool-augmented large language models (LLMs) leverages translation of natural language (NL) statements into First-Order Logic (FOL) and external theorem provers. However, the correctness of FOL statements, comprising operators and text, often go unverified due to the lack of a reliable evaluation metric for comparing generated and ground-truth FOLs. In this paper, we conduct a comprehensive study on the sensitivity of existing metrics—NL, FOL, and graph-based— and their alignment with LLM as a judge on FOL evaluation to measure robustness. We introduce operator and text-based perturbations to ground-truth FOL statements to assess metric sensitivity. We then evaluate metric robustness by comparing them against LLMs judgement. Our empirical findings highlight a clear oversensitivity in the n-gram metric BLEU for text perturbations. The operator perturbation affects the semantic graph metric Smatch++ for structural changes, and the FOL metric for specific operator changes. We observe a closer alignment between BertScore and LLM judgement, proving the importance of semantic evaluation. Additionally, we show that combining metrics enhances both robustness and sensitivity compared to using individual metrics.

BibTeX
@inproceedings{emnlp2025_assessingthesens,
  title = {Assessing the Sensitivity and Alignment of FOL Closeness Metrics},
  author = {Ramya Keerthy Thatikonda and Wray Buntine and Ehsan Shareghi},
  booktitle = {EMNLP 2025},
  year = {2025}
}
Assessing the Sensitivity and Alignment of FOL Closeness Metrics · EMNLP 2025