EMNLP 2022finding46 citations

Linguistic Rules-Based Corpus Generation for Native Chinese Grammatical Error Correction

Shirong Ma, Yinghui Li, Rongyi Sun, Qingyu Zhou, Shulin Huang, Ding Zhang, Li Yangning, Ruiyang Liu

Abstract

Chinese Grammatical Error Correction (CGEC) is both a challenging NLP task and a common application in human daily life. Recently, many data-driven approaches are proposed for the development of CGEC research. However, there are two major limitations in the CGEC field: First, the lack of high-quality annotated training corpora prevents the performance of existing CGEC models from being significantly improved. Second, the grammatical errors in widely used test sets are not made by native Chinese speakers, resulting in a significant gap between the CGEC models and the real application. In this paper, we propose a linguistic rules-based approach to construct large-scale CGEC training corpora with automatically generated grammatical errors. Additionally, we present a challenging CGEC benchmark derived entirely from errors made by native Chinese speakers in real-world scenarios. Extensive experiments and detailed analyses not only demonstrate that the training data constructed by our method effectively improves the performance of CGEC models, but also reflect that our benchmark is an excellent resource for further development of the CGEC field.

BibTeX
@inproceedings{ma-etal-2022-linguistic,
    title = "Linguistic Rules-Based Corpus Generation for Native {C}hinese Grammatical Error Correction",
    author = "Ma, Shirong  and
      Li, Yinghui  and
      Sun, Rongyi  and
      Zhou, Qingyu  and
      Huang, Shulin  and
      Zhang, Ding  and
      Yangning, Li  and
      Liu, Ruiyang  and
      Li, Zhongli  and
      Cao, Yunbo  and
      Zheng, Haitao  and
      Shen, Ying",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Findings of the Association for Computational Linguistics: EMNLP 2022",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.findings-emnlp.40/",
    doi = "10.18653/v1/2022.findings-emnlp.40",
    pages = "576--589"
}
Linguistic Rules-Based Corpus Generation for Native Chinese Grammatical Error Correction · EMNLP 2022