2025
JudgeBench: A Benchmark for Evaluating LLM-Based Judges
ICLR 2025poster
LLM-based judges have emerged as a scalable alternative to human evaluation and are increasingly used to assess, compare, and improve models. However, the reliability of LLM-based judges themselves is rarely scrutinized. As LLMs become more advanced, their responses grow more sophisticated, requirin…