Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model
Ahmet Üstün, Viraat Aryabumi, Zheng Yong, Wei-Yin Ko, Daniel D’souza, Gbemileke Onilude, Neel Bhandari, Shivalika Singh
Abstract
Recent breakthroughs in large language models (LLMs) have centered around a handful of data-rich languages. What does it take to broaden access to breakthroughs beyond first-class citizen languages? Our work introduces Aya, a massively multilingual generative language model that follows instructions in 101 languages of which over 50% are considered as lower-resourced. Aya outperforms mT0 and BLOOMZ on the majority of tasks while covering double the number of languages. We introduce extensive new evaluation suites that broaden the state-of-art for multilingual eval across 99 languages —— including discriminative and generative tasks, human evaluation, and simulated win rates that cover both held-out tasks and in-distribution performance. Furthermore, we conduct detailed investigations on the optimal finetuning mixture composition, data pruning, as well as the toxicity, bias, and safety of our models.
BibTeX
@inproceedings{ustun-etal-2024-aya,
title = "Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model",
author = {{\"U}st{\"u}n, Ahmet and
Aryabumi, Viraat and
Yong, Zheng and
Ko, Wei-Yin and
D{'}souza, Daniel and
Onilude, Gbemileke and
Bhandari, Neel and
Singh, Shivalika and
Ooi, Hui-Lee and
Kayid, Amr and
Vargus, Freddie and
Blunsom, Phil and
Longpre, Shayne and
Muennighoff, Niklas and
Fadaee, Marzieh and
Kreutzer, Julia and
Hooker, Sara},
editor = "Ku, Lun-Wei and
Martins, Andre and
Srikumar, Vivek",
booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = aug,
year = "2024",
address = "Bangkok, Thailand",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.acl-long.845/",
doi = "10.18653/v1/2024.acl-long.845",
pages = "15894--15939"
}