EMNLP 2024main35 citations

A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations

Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan

Abstract

Large Language Models (LLMs) have recently gained significant attention due to their remarkable capabilities in performing diverse tasks across various domains. However, a thorough evaluation of these models is crucial before deploying them in real-world applications to ensure they produce reliable performance. Despite the well-established importance of evaluating LLMs in the community, the complexity of the evaluation process has led to varied evaluation setups, causing inconsistencies in findings and interpretations. To address this, we systematically review the primary challenges and limitations causing these inconsistencies and unreliable evaluations in various steps of LLM evaluation. Based on our critical review, we present our perspectives and recommendations to ensure LLM evaluations are reproducible, reliable, and robust.

BibTeX
@inproceedings{laskar-etal-2024-systematic,
    title = "A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations",
    author = "Laskar, Md Tahmid Rahman  and
      Alqahtani, Sawsan  and
      Bari, M Saiful  and
      Rahman, Mizanur  and
      Khan, Mohammad Abdullah Matin  and
      Khan, Haidar  and
      Jahan, Israt  and
      Bhuiyan, Amran  and
      Tan, Chee Wei  and
      Parvez, Md Rizwan  and
      Hoque, Enamul  and
      Joty, Shafiq  and
      Huang, Jimmy",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.764/",
    doi = "10.18653/v1/2024.emnlp-main.764",
    pages = "13785--13816"
}
A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and Recommendations · EMNLP 2024