2021
Automatic Construction of Evaluation Suites for Natural Language Generation Datasets
NeurIPS 2021poster
Machine learning approaches applied to NLP are often evaluated by summarizing their performance in a single number, for example accuracy. Since most test sets are constructed as an i.i.d. sample from the overall data, this approach overly simplifies the complexity of language and encourages overfitt…