← Search

Tan Hongzhi

1 accepted papers

2025

Sample-Efficient Human Evaluation of Large Language Models via Maximum Discrepancy Competition

ACL 2025long

The past years have witnessed a proliferation of large language models (LLMs). Yet, reliable evaluation of LLMs is challenging due to the inaccuracy of standard metrics in human perception of text quality and the inefficiency in sampling informative test examples for human evaluation. This paper pre…