2024
How Reliable Are Automatic Evaluation Methods for Instruction-Tuned LLMs?
EMNLP 2024finding
Work on instruction-tuned Large Language Models (LLMs) has used automatic methods based on text overlap and LLM judgments as cost-effective alternatives to human evaluation. In this paper, we perform a meta-evaluation of such methods and assess their reliability across a broad range of tasks. In eva…