2022
Detecting Label Errors by Using Pre-Trained Language Models
EMNLP 2022main
We show that large pre-trained language models are inherently highly capable of identifying label errors in natural language datasets: simply examining out-of-sample data points in descending order of fine-tuned task loss significantly outperforms more complex error-detection mechanisms proposed in…