ACL 2024findings10 citations

Evaluating Large Language Models on Wikipedia-Style Survey Generation

Fan Gao, Hang Jiang, Rui Yang, Qingcheng Zeng, Jinghui Lu, Moritz Blum, Tianwei She, Yuang Jiang

Abstract

Educational materials such as survey articles in specialized fields like computer science traditionally require tremendous expert inputs and are therefore expensive to create and update. Recently, Large Language Models (LLMs) have achieved significant success across various general tasks. However, their effectiveness and limitations in the education domain are yet to be fully explored. In this work, we examine the proficiency of LLMs in generating succinct survey articles specific to the niche field of NLP in computer science, focusing on a curated list of 99 topics. Automated benchmarks reveal that GPT-4 surpasses its predecessors, inluding GPT-3.5, PaLM2, and LLaMa2 by margins ranging from 2% to 20% in comparison to the established ground truth. We compare both human and GPT-based evaluation scores and provide in-depth analysis. While our findings suggest that GPT-created surveys are more contemporary and accessible than human-authored ones, certain limitations were observed. Notably, GPT-4, despite often delivering outstanding content, occasionally exhibited lapses like missing details or factual errors. At last, we compared the rating behavior between humans and GPT-4 and found systematic bias in using GPT evaluation.

BibTeX
@inproceedings{gao-etal-2024-evaluating-large,
    title = "Evaluating Large Language Models on {W}ikipedia-Style Survey Generation",
    author = "Gao, Fan  and
      Jiang, Hang  and
      Yang, Rui  and
      Zeng, Qingcheng  and
      Lu, Jinghui  and
      Blum, Moritz  and
      She, Tianwei  and
      Jiang, Yuang  and
      Li, Irene",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-acl.321/",
    doi = "10.18653/v1/2024.findings-acl.321",
    pages = "5405--5418"
}
Evaluating Large Language Models on Wikipedia-Style Survey Generation · ACL 2024