ACL 2024long8 citations

One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation

Tejpalsingh Siledar, Swaroop Nath, Sankara Muddu, Rupasai Rangaraju, Swaprava Nath, Pushpak Bhattacharyya, Suman Banerjee, Amey Patil

Abstract

Evaluation of opinion summaries using conventional reference-based metrics often fails to provide a comprehensive assessment and exhibits limited correlation with human judgments. While Large Language Models (LLMs) have shown promise as reference-free metrics for NLG evaluation, their potential remains unexplored for opinion summary evaluation. Furthermore, the absence of sufficient opinion summary evaluation datasets hinders progress in this area. In response, we introduce the SUMMEVAL-OP dataset, encompassing 7 dimensions crucial to the evaluation of opinion summaries: fluency, coherence, relevance, faithfulness, aspect coverage, sentiment consistency, and specificity. We propose OP-I-PROMPT, a dimension-independent prompt, along with OP-PROMPTS, a dimension-dependent set of prompts for opinion summary evaluation. Our experiments demonstrate that OP-I-PROMPT emerges as a good alternative for evaluating opinion summaries, achieving an average Spearman correlation of 0.70 with human judgments, surpassing prior methodologies. Remarkably, we are the first to explore the efficacy of LLMs as evaluators, both on closed-source and open-source models, in the opinion summary evaluation domain.

BibTeX
@inproceedings{siledar-etal-2024-one,
    title = "One Prompt To Rule Them All: {LLM}s for Opinion Summary Evaluation",
    author = "Siledar, Tejpalsingh  and
      Nath, Swaroop  and
      Muddu, Sankara  and
      Rangaraju, Rupasai  and
      Nath, Swaprava  and
      Bhattacharyya, Pushpak  and
      Banerjee, Suman  and
      Patil, Amey  and
      Singh, Sudhanshu  and
      Chelliah, Muthusamy  and
      Garera, Nikesh",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.655/",
    doi = "10.18653/v1/2024.acl-long.655",
    pages = "12119--12134"
}
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation · ACL 2024