ACL 2022long136 citations
XLM-E: Cross-lingual Language Model Pre-training via ELECTRA
Zewen Chi, Shaohan Huang, Li Dong, Shuming Ma, Bo Zheng, Saksham Singhal, Payal Bajaj, Xia Song
Abstract
In this paper, we introduce ELECTRA-style tasks to cross-lingual language model pre-training. Specifically, we present two pre-training tasks, namely multilingual replaced token detection, and translation replaced token detection. Besides, we pretrain the model, named as XLM-E, on both multilingual and parallel corpora. Our model outperforms the baseline models on various cross-lingual understanding tasks with much less computation cost. Moreover, analysis shows that XLM-E tends to obtain better cross-lingual transferability.
BibTeX
@inproceedings{chi-etal-2022-xlm,
title = "{XLM}-{E}: Cross-lingual Language Model Pre-training via {ELECTRA}",
author = "Chi, Zewen and
Huang, Shaohan and
Dong, Li and
Ma, Shuming and
Zheng, Bo and
Singhal, Saksham and
Bajaj, Payal and
Song, Xia and
Mao, Xian-Ling and
Huang, Heyan and
Wei, Furu",
editor = "Muresan, Smaranda and
Nakov, Preslav and
Villavicencio, Aline",
booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = may,
year = "2022",
address = "Dublin, Ireland",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2022.acl-long.427/",
doi = "10.18653/v1/2022.acl-long.427",
pages = "6170--6182"
}