NAACL 2025long0 citations

Evaluating Evidence Attribution in Generated Fact Checking Explanations

Rui Xing, Timothy Baldwin, Jey Han Lau

Abstract

Automated fact-checking systems often struggle with trustworthiness, as their generated explanations can include hallucinations. In this work, we explore evidence attribution for fact-checking explanation generation. We introduce a novel evaluation protocol, citation masking and recovery, to assess attribution quality in generated explanations. We implement our protocol using both human annotators and automatic annotators and found that LLM annotation correlates with human annotation, suggesting that attribution assessment can be automated. Finally, our experiments reveal that: (1) the best-performing LLMs still generate explanations that are not always accurate in their attribution; and (2) human-curated evidence is essential for generating better explanations.

BibTeX
@inproceedings{xing-etal-2025-evaluating,
    title = "Evaluating Evidence Attribution in Generated Fact Checking Explanations",
    author = "Xing, Rui  and
      Baldwin, Timothy  and
      Lau, Jey Han",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.282/",
    pages = "5475--5496",
    ISBN = "979-8-89176-189-6"
}
Evaluating Evidence Attribution in Generated Fact Checking Explanations · NAACL 2025