Language Adaptation of Large Language Models: An Empirical Study on LLaMA2
Shumin Wang, Yuexiang Xie, Bolin Ding, Jinyang Gao, Yanyong Zhang
Abstract
There has been a surge of interest regarding language adaptation of Large Language Models (LLMs) to enhance the processing of texts in low-resource languages. While traditional language models have seen extensive research on language transfer, modern LLMs still necessitate further explorations in language adaptation. In this paper, we present a systematic review of the language adaptation process for LLMs, including vocabulary expansion, continued pre-training, and instruction fine-tuning, which focuses on empirical studies conducted on LLaMA2 and discussions on various settings affecting the model’s capabilities. This study provides helpful insights covering the entire language adaptation process, and highlights the compatibility and interactions between different steps, offering researchers a practical guidebook to facilitate the effective adaptation of LLMs across different languages.
BibTeX
@inproceedings{wang-etal-2025-language,
title = "Language Adaptation of Large Language Models: An Empirical Study on {LL}a{MA}2",
author = "Wang, Shumin and
Xie, Yuexiang and
Ding, Bolin and
Gao, Jinyang and
Zhang, Yanyong",
editor = "Rambow, Owen and
Wanner, Leo and
Apidianaki, Marianna and
Al-Khalifa, Hend and
Eugenio, Barbara Di and
Schockaert, Steven",
booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
month = jan,
year = "2025",
address = "Abu Dhabi, UAE",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.coling-main.480/",
pages = "7195--7208"
}