← Search

S{\'e}b Arnold

1 accepted papers

2025

Graders Should Cheat: Privileged Information Enables Expert-Level Automated Evaluations

EMNLP 2025

Auto-evaluating language models (LMs), *i.e*., using a grader LM to evaluate the candidate LM, is an appealing way to accelerate the evaluation process and the cost associated with it. But this presents a paradox: how can we trust the grader LM, which is presumably weaker than the candidate LM, to a

Cited by 0SourcePDFScholar