← Search

William Huang

4 accepted papers

2021

Comparing Test Sets with Item Response Theory

ACL 2021long

Recent years have seen numerous NLP datasets introduced to evaluate the performance of fine-tuned models on natural language understanding tasks. Recent results from large pretrained models, though, show that many of these datasets are largely saturated and unlikely to be able to detect further prog…

Cited by 45SourcePDFScholar
2021

Does Putting a Linguist in the Loop Improve NLU Data Collection?

EMNLP 2021finding

Many crowdsourced NLP datasets contain systematic artifacts that are identified only after data collection is complete. Earlier identification of these issues should make it easier to create high-quality training and evaluation data. We attempt this by evaluating protocols in which expert linguists…

Cited by 46SourcePDFScholar