2021
Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks
NeurIPS 2021poster
We identify label errors in the test sets of 10 of the most commonly-used computer vision, natural language, and audio datasets, and subsequently study the potential for these label errors to affect benchmark results. Errors in test sets are numerous and widespread: we estimate an average of at lea…