ACL 2025long0 citations

Improving Model Factuality with Fine-grained Critique-based Evaluator

Yiqing Xie, Wenxuan Zhou, Pradyot Prakash, Di Jin, Yuning Mao, Quintin Fettes, Arya Talebzadeh, Sinong Wang

Abstract

Factuality evaluation aims to detect factual errors produced by language models (LMs) and hence guide the development of more factual models. Towards this goal, we train a factuality evaluator, FenCE, that provides LM generators with claim-level factuality feedback. In particular, we train FenCE to (1) generate textual critiques along with scores and (2) make claim-level judgment based on diverse source documents obtained by various tools, via data augmentation on a combination of public judgment datasets. We then present a framework that leverages FenCE to improve the factuality of LM generators by constructing training data. Specifically, we generate a set of candidate responses, ask FenCE to revise and score each response without introducing lesser-known facts, and train the generator by preferring highly scored revised responses. Experiments show that our data augmentation methods improve the evaluator’s accuracy by 2.9% on LLM-AggreFact. With FenCE, we improve Llama2-7B-chat/Llama3-8B-chat’s factuality rate by 16.86%/14.45% on FActScore, outperforming state-of-the-art factuality finetuning methods by 8.83%/6.96%.

BibTeX
@inproceedings{xie-etal-2025-improving,
    title = "Improving Model Factuality with Fine-grained Critique-based Evaluator",
    author = "Xie, Yiqing  and
      Zhou, Wenxuan  and
      Prakash, Pradyot  and
      Jin, Di  and
      Mao, Yuning  and
      Fettes, Quintin  and
      Talebzadeh, Arya  and
      Wang, Sinong  and
      Fang, Han  and
      Rose, Carolyn  and
      Fried, Daniel  and
      Zhang, Hejia",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.400/",
    doi = "10.18653/v1/2025.acl-long.400",
    pages = "8140--8155",
    ISBN = "979-8-89176-251-0"
}
Improving Model Factuality with Fine-grained Critique-based Evaluator · ACL 2025