2025
Position: Don't Use the CLT in LLM Evals With Fewer Than a Few Hundred Datapoints
ICML 2025spotlight
Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessment. Currently, when such statistical measures are reported, they typically rely on the Central Limit Theorem (CLT). In…