2026
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
ICLR 2026poster
The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundamental types of inconsistencies: (1) \textit{Score-Comparison Inconsistency}, where lower-rated responses outperform high…