NAACL 2025long7 citations

IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models

David Ifeoluwa Adelani, Jessica Ojo, Israel Abebe Azime, Jian Yun Zhuang, Jesujoba Oluwadara Alabi, Xuanli He, Millicent Ochieng, Sara Hooker

Abstract

Despite the widespread adoption of Large language models (LLMs), their remarkable capabilities remain limited to a few high-resource languages. Additionally, many low-resource languages (e.g. African languages) are often evaluated only on basic text classification tasks due to the lack of appropriate or comprehensive benchmarks outside of high-resource languages. In this paper, we introduce IrokoBench—a human-translated benchmark dataset for 17 typologically-diverse low-resource African languages covering three tasks: natural language inference(AfriXNLI), mathematical reasoning(AfriMGSM), and multi-choice knowledge-based QA(AfriMMLU). We use IrokoBench to evaluate zero-shot, few-shot, and translate-test settings(where test sets are translated into English) across 10 open and four proprietary LLMs. Our evaluation reveals a significant performance gap between high-resource languages (such as English and French) and low-resource African languages. We observe a significant performance gap between open and proprietary models, with the highest performing open model, Gemma 2 27B only at 63% of the best-performing proprietary model GPT-4o performance. Machine translating the test set to English before evaluation helped to close the gap for larger models that are English-centric, like Gemma 2 27B and LLaMa 3.1 70B. These findings suggest that more efforts are needed to develop and adapt LLMs for African languages.

BibTeX
@inproceedings{adelani-etal-2025-irokobench,
    title = "{I}roko{B}ench: A New Benchmark for {A}frican Languages in the Age of Large Language Models",
    author = "Adelani, David Ifeoluwa  and
      Ojo, Jessica  and
      Azime, Israel Abebe  and
      Zhuang, Jian Yun  and
      Alabi, Jesujoba Oluwadara  and
      He, Xuanli  and
      Ochieng, Millicent  and
      Hooker, Sara  and
      Bukula, Andiswa  and
      Lee, En-Shiun Annie  and
      Chukwuneke, Chiamaka Ijeoma  and
      Buzaaba, Happy  and
      Sibanda, Blessing Kudzaishe  and
      Kalipe, Godson Koffi  and
      Mukiibi, Jonathan  and
      Kabongo Kabenamualu, Salomon  and
      Yuehgoh, Foutse  and
      Setaka, Mmasibidi  and
      Ndolela, Lolwethu  and
      Odu, Nkiruka  and
      Mabuya, Rooweither  and
      Osei, Salomey  and
      Muhammad, Shamsuddeen Hassan  and
      Samb, Sokhar  and
      Guge, Tadesse Kebede  and
      Sherman, Tombekai Vangoni  and
      Stenetorp, Pontus",
    editor = "Chiruzzo, Luis  and
      Ritter, Alan  and
      Wang, Lu",
    booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = apr,
    year = "2025",
    address = "Albuquerque, New Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.naacl-long.139/",
    pages = "2732--2757",
    ISBN = "979-8-89176-189-6"
}