← Search

Huanxin Sheng

1 accepted papers

2025

Analyzing Uncertainty of LLM-as-a-Judge: Interval Evaluations with Conformal Prediction

EMNLP 2025

LLM-as-a-judge has become a promising paradigm for using large language models (LLMs) to evaluate natural language generation (NLG), but the uncertainty of its evaluation remains underexplored. This lack of reliability may limit its deployment in many applications. This work presents the first frame

Cited by 0SourcePDFScholar