← Search

Gonghan Xu

1 accepted papers

2024

Bayesian Calibration of Win Rate Estimation with LLM Evaluators

EMNLP 2024main

Recent advances in large language models (LLMs) show the potential of using LLMs as evaluators for assessing the quality of text generations from LLMs. However, applying LLM evaluators naively to compare different systems can lead to unreliable results due to the inaccuracy and intrinsic bias of LLM…