← Search

Gaonan Chen

1 accepted papers

2025

UBench: Benchmarking Uncertainty in Large Language Models with Multiple Choice Questions

ACL 2025finding

Despite recent progress in systematic evaluation frameworks, benchmarking the uncertainty of large language models (LLMs) remains a highly challenging task. Existing methods for benchmarking the uncertainty of LLMs face three key challenges: the need for internal model access, additional training, o…