2024
Bayesian Calibration of Win Rate Estimation with LLM Evaluators
EMNLP 2024main
Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evaluators naively to compare different systems can lead to unreliable results due to the inaccuracy and intrinsic bias of LLM…