EMNLP 2024main3 citations

SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages

Holy Lovenia, Rahmad Mahendra, Salsabil Maulana Akbar, Lester James Validad Miranda, Jennifer Santoso, Elyanah Aco, Akhdan Fadhilah, Jonibek Mansurov

Abstract

Southeast Asia (SEA) is a region rich in linguistic diversity and cultural variety, with over 1,300 indigenous languages and a population of 671 million people. However, prevailing AI models suffer from a significant lack of representation of texts, images, and audio datasets from SEA, compromising the quality of AI models for SEA languages. Evaluating models for SEA languages is challenging due to the scarcity of high-quality datasets, compounded by the dominance of English training data, raising concerns about potential cultural misrepresentation. To address these challenges, through a collaborative movement, we introduce SEACrowd, a comprehensive resource center that fills the resource gap by providing standardized corpora in nearly 1,000 SEA languages across three modalities. Through our SEACrowd benchmarks, we assess the quality of AI models on 36 indigenous languages across 13 tasks, offering valuable insights into the current AI landscape in SEA. Furthermore, we propose strategies to facilitate greater AI advancements, maximizing potential utility and resource equity for the future of AI in Southeast Asia.

BibTeX
@inproceedings{lovenia-etal-2024-seacrowd,
    title = "{SEAC}rowd: A Multilingual Multimodal Data Hub and Benchmark Suite for {S}outheast {A}sian Languages",
    author = {Lovenia, Holy  and
      Mahendra, Rahmad  and
      Akbar, Salsabil Maulana  and
      Miranda, Lester James Validad  and
      Santoso, Jennifer  and
      Aco, Elyanah  and
      Fadhilah, Akhdan  and
      Mansurov, Jonibek  and
      Imperial, Joseph Marvin  and
      Kampman, Onno P.  and
      Moniz, Joel Ruben Antony  and
      Habibi, Muhammad Ravi Shulthan  and
      Hudi, Frederikus  and
      Montalan, Railey  and
      Hadiwijaya, Ryan Ignatius  and
      Lopo, Joanito Agili  and
      Nixon, William  and
      Karlsson, B{\"o}rje F.  and
      Jaya, James  and
      Diandaru, Ryandito  and
      Gao, Yuze  and
      Irawan, Patrick Amadeus  and
      Wang, Bin  and
      Cruz, Jan Christian Blaise  and
      Whitehouse, Chenxi  and
      Parmonangan, Ivan Halim  and
      Khelli, Maria  and
      Zhang, Wenyu  and
      Susanto, Lucky  and
      Ryanda, Reynard Adha  and
      Hermawan, Sonny Lazuardi  and
      Velasco, Dan John  and
      Kautsar, Muhammad Dehan Al  and
      Hendria, Willy Fitra  and
      Moslem, Yasmin  and
      Flynn, Noah  and
      Adilazuarda, Muhammad Farid  and
      Li, Haochen  and
      Lee, Johanes  and
      Damanhuri, R.  and
      Sun, Shuo  and
      Qorib, Muhammad Reza  and
      Djanibekov, Amirbek  and
      Leong, Wei Qi  and
      Do, Quyet V.  and
      Muennighoff, Niklas  and
      Pansuwan, Tanrada  and
      Putra, Ilham Firdausi  and
      Xu, Yan  and
      Chia, Tai Ngee  and
      Purwarianti, Ayu  and
      Ruder, Sebastian  and
      Tjhi, William Chandra  and
      Limkonchotiwat, Peerat  and
      Aji, Alham Fikri  and
      Keh, Sedrick  and
      Winata, Genta Indra  and
      Zhang, Ruochen  and
      Koto, Fajri  and
      Yong, Zheng Xin  and
      Cahyawijaya, Samuel},
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.296/",
    doi = "10.18653/v1/2024.emnlp-main.296",
    pages = "5155--5203"
}
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages · EMNLP 2024