← Search

Peiyu Li

1 accepted papers

2026

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

ICML 2026spotlight

Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale. Existing evaluation protocols rely on average accuracy over fixed item sets, treating all items as equally informative despite substanti…

Cited by 0SourceScholar