2025
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
NeurIPS 2025poster
Evaluating natural language generation (NLG) systems remains a core challenge, further complicated by the rise of general-purpose large language models (LLMs). Recently, large language models as judges (LLJs) have emerged as a scalable, cost-effective alternative to traditional metrics, but their va…