2025
Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations
EMNLP 2025
Auto-evaluating language models (LMs), *i.e*., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to a