EMNLP 2024main2 citations

Concept Space Alignment in Multilingual LLMs

Qiwei Peng, Anders Søgaard

Abstract

Multilingual large language models (LLMs) seem to generalize somewhat across languages. We hypothesize this is a result of implicit vector space alignment. Evaluating such alignment, we see that larger models exhibit very high-quality linear alignments between corresponding concepts in different languages. Our experiments show that multilingual LLMs suffer from two familiar weaknesses: generalization works best for languages with similar typology, and for abstract concepts. For some models, e.g., the Llama-2 family of models, prompt-based embeddings align better than word embeddings, but the projections are less linear – an observation that holds across almost all model families, indicating that some of the implicitly learned alignments are broken somewhat by prompt-based methods.

BibTeX
@inproceedings{peng-sogaard-2024-concept,
    title = "Concept Space Alignment in Multilingual {LLM}s",
    author = "Peng, Qiwei  and
      S{\o}gaard, Anders",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.315/",
    doi = "10.18653/v1/2024.emnlp-main.315",
    pages = "5511--5526"
}
Concept Space Alignment in Multilingual LLMs · EMNLP 2024