SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James Validad Miranda, Jennifer Santoso, Elyanah Aco, Akhdan Fadhilah, Jonibek Mansurov
Abstract
Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA, compromising the quality of AI models for SEA languages. Evaluating models for SEA languages is challenging due to the scarcity of high-quality datasets, compounded by the dominance of English training data, raising concerns about potential cultural misrepresentation. To address these challenges, through a collaborative movement, we introduce SEACrowd, a comprehensive resource center that fills the resource gap by providing standardized corpora in nearly 1,000 SEA languages across three modalities. Through our SEACrowd benchmarks, we assess the quality of AI models on 36 indigenous languages across 13 tasks, offering valuable insights into the current AI landscape in SEA. Furthermore, we propose strategies to facilitate greater AI advancements, maximizing potential utility and resource equity for the future of AI in Southeast Asia.
BibTeX
@inproceedings{lovenia-etal-2024-seacrowd,
title = "{SEAC}rowd: A Multilingual Multimodal Data Hub and Benchmark Suite for {S}outheast {A}sian Languages",
author = {Lovenia, Holy and
Mahendra, Rahmad and
Akbar, Salsabil Maulana and
Miranda, Lester James Validad and
Santoso, Jennifer and
Aco, Elyanah and
Fadhilah, Akhdan and
Mansurov, Jonibek and
Imperial, Joseph Marvin and
Kampman, Onno P. and
Moniz, Joel Ruben Antony and
Habibi, Muhammad Ravi Shulthan and
Hudi, Frederikus and
Montalan, Railey and
Hadiwijaya, Ryan Ignatius and
Lopo, Joanito Agili and
Nixon, William and
Karlsson, B{\"o}rje F. and
Jaya, James and
Diandaru, Ryandito and
Gao, Yuze and
Irawan, Patrick Amadeus and
Wang, Bin and
Cruz, Jan Christian Blaise and
Whitehouse, Chenxi and
Parmonangan, Ivan Halim and
Khelli, Maria and
Zhang, Wenyu and
Susanto, Lucky and
Ryanda, Reynard Adha and
Hermawan, Sonny Lazuardi and
Velasco, Dan John and
Kautsar, Muhammad Dehan Al and
Hendria, Willy Fitra and
Moslem, Yasmin and
Flynn, Noah and
Adilazuarda, Muhammad Farid and
Li, Haochen and
Lee, Johanes and
Damanhuri, R. and
Sun, Shuo and
Qorib, Muhammad Reza and
Djanibekov, Amirbek and
Leong, Wei Qi and
Do, Quyet V. and
Muennighoff, Niklas and
Pansuwan, Tanrada and
Putra, Ilham Firdausi and
Xu, Yan and
Chia, Tai Ngee and
Purwarianti, Ayu and
Ruder, Sebastian and
Tjhi, William Chandra and
Limkonchotiwat, Peerat and
Aji, Alham Fikri and
Keh, Sedrick and
Winata, Genta Indra and
Zhang, Ruochen and
Koto, Fajri and
Yong, Zheng Xin and
Cahyawijaya, Samuel},
editor = "Al-Onaizan, Yaser and
Bansal, Mohit and
Chen, Yun-Nung",
booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
month = nov,
year = "2024",
address = "Miami, Florida, USA",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2024.emnlp-main.296/",
doi = "10.18653/v1/2024.emnlp-main.296",
pages = "5155--5203"
}