2026
A Judge-Aware Ranking Framework for Evaluating Large Language Models without Ground Truth
ICML 2026poster
Evaluating large language models (LLMs) on open-ended tasks without ground-truth labels is increasingly done via the LLM-as-a-judge paradigm. A critical but under-modeled issue is that judge LLMs differ substantially in reliability; treating all judges equally can yield biased leaderboards and misle…