2025
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
ICML 2025poster
Evaluating Large Language Models (LLMs) often requires costly human annotations. To address this, LLM-based judges have been proposed, which compare the outputs of two LLMs enabling the ranking of models without human intervention. While several approaches have been proposed, many confounding facto…