COLING 2025main0 citations

PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation

Hour Kaing, Raj Dabre, Haiyue Song, Van-Hien Tran, Hideki Tanaka, Masao Utiyama

Abstract

This work introduces PrahokBART, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus quality and addressing the linguistic issues of Khmer, which are ignored in existing multilingual models, by incorporating linguistic components such as word segmentation and normalization. We evaluate PrahokBART on three generative tasks: machine translation, text summarization, and headline generation, where our results demonstrate that it outperforms mBART50, a strong multilingual pre-trained model. Additionally, our analysis provides insights into the impact of each linguistic module and evaluates how effectively our model handles space during text generation, which is crucial for the naturalness of texts in Khmer.

BibTeX
@inproceedings{kaing-etal-2025-prahokbart,
    title = "{P}rahok{BART}: A Pre-trained Sequence-to-Sequence Model for {K}hmer Natural Language Generation",
    author = "Kaing, Hour  and
      Dabre, Raj  and
      Song, Haiyue  and
      Tran, Van-Hien  and
      Tanaka, Hideki  and
      Utiyama, Masao",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.87/",
    pages = "1309--1322"
}
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation · COLING 2025