COLING 2025industry0 citations

Leveraging Multilingual Models for Robust Grammatical Error Correction Across Low-Resource Languages

Divesh Ramesh Kubal, Apurva Shrikant Nagvenkar

Abstract

Grammatical Error Correction (GEC) is a crucial task in Natural Language Processing (NLP) aimed at improving the quality of user-generated content, particularly for non-native speakers. This paper introduces a novel end-to-end architecture utilizing the M2M100 multilingual transformer model to build a unified GEC system, with a focus on low-resource languages. A synthetic data generation pipeline is proposed, tailored to address language-specific error categories. The system has been implemented for the Spanish language, showing promising results based on evaluations conducted by linguists with expertise in Spanish. Additionally, we present a user analysis that tracks user interactions, revealing an acceptance rate of 88.2%, as reflected by the actions performed by users.

BibTeX
@inproceedings{kubal-nagvenkar-2025-leveraging,
    title = "Leveraging Multilingual Models for Robust Grammatical Error Correction Across Low-Resource Languages",
    author = "Kubal, Divesh Ramesh  and
      Nagvenkar, Apurva Shrikant",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven  and
      Darwish, Kareem  and
      Agarwal, Apoorv",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics: Industry Track",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-industry.43/",
    pages = "505--510"
}
Leveraging Multilingual Models for Robust Grammatical Error Correction Across Low-Resource Languages · COLING 2025