← Search

Raluca Popa

1 accepted papers

2025

JudgeBench: A Benchmark for Evaluating LLM-Based Judges

ICLR 2025poster

LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requirin…