2025
Position: AI Evaluation Should Learn from How We Test Humans
ICML 2025poster
As AI systems continue to evolve, their rigorous evaluation becomes crucial for their development and deployment. Researchers have constructed various large-scale benchmarks to determine their capabilities, typically against a gold-standard test set and report metrics averaged across all items. Howe…