ACL 2025long0 citations

Can LLMs Evaluate Complex Attribution in QA? Automatic Benchmarking using Knowledge Graphs

Nan Hu, Jiaoyan Chen, Yike Wu, Guilin Qi, Hongru Wang, Sheng Bi, Yongrui Chen, Tongtong Wu

Abstract

Attributed Question Answering (AQA) has attracted wide attention, but there are still several limitations in evaluating the attributions, including lacking fine-grained attribution categories, relying on manual annotations, and failing to compare attributions with only subtle differences. To bridge these gaps, we introduce Complex Attributed Question Answering (CAQA), a large-scale benchmark containing comprehensive attribution categories, automatically generated using Knowledge Graphs (KGs), and complex attribution scenarios. We have conducted extensive experiments to verify the effectiveness of CAQA, including the benchmarking of 25 automatic evaluators, their comparison with human evaluators, the testing of LLM evaluators fine-tuned by CAQA and so on. These experiments also lead to a series of important findings that can benefit the future research of AQA.

BibTeX
@inproceedings{hu-etal-2025-llms,
    title = "Can {LLM}s Evaluate Complex Attribution in {QA}? Automatic Benchmarking using Knowledge Graphs",
    author = "Hu, Nan  and
      Chen, Jiaoyan  and
      Wu, Yike  and
      Qi, Guilin  and
      Wang, Hongru  and
      Bi, Sheng  and
      Chen, Yongrui  and
      Wu, Tongtong  and
      Pan, Jeff Z.",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.837/",
    doi = "10.18653/v1/2025.acl-long.837",
    pages = "17096--17118",
    ISBN = "979-8-89176-251-0"
}