EMNLP 2022main28 citations

Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing

Abbas Ghaddar, Yimeng Wu, Sunyam Bagga, Ahmad Rashid, Khalil Bibi, Mehdi Rezagholizadeh, Chao Xing, Yasheng Wang

Abstract

There is a growing body of work in recent years to develop pre-trained language models (PLMs) for the Arabic language. This work addresses two major problems in existing Arabic PLMs that limit the progress of the Arabic NLU and NLG fields. First, existing Arabic PLMs are not well-explored and their pre-training can be improved significantly using a more methodical approach. Second, there is a lack of systematic and reproducible evaluation of these models in the literature. We revisit both the pre-training and evaluation of Arabic PLMs. In terms of pre-training, we explore the impact of the quality of the pretraining data, the size of the model, and the incorporation of character-level information on Arabic PLM. As a result, we release three new Arabic BERT-style models ( JABER, Char-JABER, and SABER), and two T5-style models (AT5S and AT5B). In terms of evaluation, we conduct a comprehensive empirical study to systematically evaluate the performance of existing state-of-the-art models on ALUE, a leaderboard-powered benchmark for Arabic NLU tasks, and on a subset of the Arabic generative tasks. We show that our models significantly outperform existing Arabic PLMs and achieve a new state-of-the-art performance on discriminative and generative Arabic NLU and NLG tasks. Our models and source code to reproduce results will be made available upon acceptance.

BibTeX
@inproceedings{ghaddar-etal-2022-revisiting,
    title = "Revisiting Pre-trained Language Models and their Evaluation for {A}rabic Natural Language Processing",
    author = "Ghaddar, Abbas  and
      Wu, Yimeng  and
      Bagga, Sunyam  and
      Rashid, Ahmad  and
      Bibi, Khalil  and
      Rezagholizadeh, Mehdi  and
      Xing, Chao  and
      Wang, Yasheng  and
      Duan, Xinyu  and
      Wang, Zhefeng  and
      Huai, Baoxing  and
      Jiang, Xin  and
      Liu, Qun  and
      Langlais, Phillippe",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.205/",
    doi = "10.18653/v1/2022.emnlp-main.205",
    pages = "3135--3151"
}
Revisiting Pre-trained Language Models and their Evaluation for Arabic Natural Language Processing · EMNLP 2022