ACL 2024findings1 citations

Towards Better Utilization of Multi-Reference Training Data for Chinese Grammatical Error Correction

Yumeng Liu, Zhenghua Li, HaoChen Jiang, Bo Zhang, Chen Li, Ji Zhang

Abstract

For the grammatical error correction (GEC) task, there usually exist multiple correction ways for an erroneous input sentence, leading to multiple references. Observing the high proportion of multi-reference instances in Chinese GEC training data, we target a systematic study on how to better utilize multi-reference training data. We propose two new approaches and a simple two-stage training strategy. We compare them against previously proposed approaches, on two Chinese training datasets, i.e., Lang-8 for second language learner texts and FCGEC-Train for native speaker texts, and three test datasets. The experiments and analyses demonstrate the effectiveness of our proposed approaches and reveal interesting insights. Our code is available at https://github.com/ymliucs/MrGEC.

BibTeX
@inproceedings{liu-etal-2024-towards-better,
    title = "Towards Better Utilization of Multi-Reference Training Data for {C}hinese Grammatical Error Correction",
    author = "Liu, Yumeng  and
      Li, Zhenghua  and
      Jiang, HaoChen  and
      Zhang, Bo  and
      Li, Chen  and
      Zhang, Ji",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2024",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.findings-acl.180/",
    doi = "10.18653/v1/2024.findings-acl.180",
    pages = "3044--3052"
}
Towards Better Utilization of Multi-Reference Training Data for Chinese Grammatical Error Correction · ACL 2024