2025
The Progress Illusion: Revisiting meta-evaluation standards of LLM evaluators
EMNLP 2025
LLM judges have gained popularity as an inexpensive and performant substitute for human evaluation. However, we observe that the meta-evaluation setting in which the reliability of these LLM evaluators is established is substantially different from their use in model development. To address this, we