2026
Position: AI Evaluations Should be Grounded on a Theory of Capability
ICML 2026poster
Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know that a reported accuracy genuinely reflects a model’s underlying performance? Al…