2026
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
AAAI 2026technical
The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of ben