EMNLP 2022main14 citations

Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition

Shuguang Chen, Leonardo Neves, Thamar Solorio

Abstract

In this work, we take the named entity recognition task in the English language as a case study and explore style transfer as a data augmentation method to increase the size and diversity of training data in low-resource scenarios. We propose a new method to effectively transform the text from a high-resource domain to a low-resource domain by changing its style-related attributes to generate synthetic data for training. Moreover, we design a constrained decoding algorithm along with a set of key ingredients for data selection to guarantee the generation of valid and coherent data. Experiments and analysis on five different domain pairs under different data regimes demonstrate that our approach can significantly improve results compared to current state-of-the-art data augmentation methods. Our approach is a practical solution to data scarcity, and we expect it to be applicable to other NLP tasks.

BibTeX
@inproceedings{chen-etal-2022-style,
    title = "Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition",
    author = "Chen, Shuguang  and
      Neves, Leonardo  and
      Solorio, Thamar",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.120/",
    doi = "10.18653/v1/2022.emnlp-main.120",
    pages = "1827--1841"
}
Style Transfer as Data Augmentation: A Case Study on Named Entity Recognition · EMNLP 2022