← Search

Michael Hardy

3 accepted papers

2026

Noise Tectonics: Measuring the Stability of AI Benchmark Ecosystems

ICML 2026poster

AI benchmark ecosystems compress rich evaluation data into aggregate leaderboard scores, but these scores contain substantial measurement noise whose sources and magnitudes remain unquantified. Without systematic methods to measure this noise and separate signal from artifact, it is unclear when ben…

Cited by 0SourceScholar
2025

Fantastic Bugs and Where to Find Them in AI Benchmarks

NeurIPS 2025poster

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of benchmark questions is not only infeasible but also a critical bottleneck for reliable evaluation. In this work, we int…

Cited by 0SourceScholar
2025

“All that Glitters”: Techniques for Evaluations with Unreliable Model and Human Annotations

NAACL 2025findings

“Gold” and “ground truth” human-mediated labels have error. This error can escape commonly reported metrics of label quality or obscure questions of accuracy, bias, fairness, and usefulness during model evaluation. This study demonstrates methods for answering such questions even in the context of v…

Cited by 0SourcePDFScholar