Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio
Abstract
Reliable multilingual evaluation is difficult, and culturally appropriate evaluation is even harder to achieve.A common practice to fill this gap is to machine-translate English evaluation sets. However, translation introduces language bias and carries over cultural and regional assumptions from the original questions – often testing knowledge irrelevant to the target audience. In this work, we highlight the extent and impact of these biases and present a multilingual evaluation framework that aims to mitigate them through improved translations and annotation practices.Through a large-scale study involving professional and community translators and annotators, we show that state-of-the-art models excel primarily by learning Western-centric concepts. Notably, we find that model rankings on the full MMLU change when evaluated on a subset of questions explicitly marked as culturally sensitive.We release Global MMLU, a multilingual extension of MMLU across 42 languages, featuring improved translation quality, expanded language coverage, and designated subsets labeled as culturally sensitive and culturally agnostic to enable a more comprehensive and equitable benchmark for evaluating language models across diverse linguistic and cultural contexts.
BibTeX
@inproceedings{singh-etal-2025-global,
title = "Global {MMLU}: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation",
author = "Singh, Shivalika and
Romanou, Angelika and
Fourrier, Cl{\'e}mentine and
Adelani, David Ifeoluwa and
Ngui, Jian Gang and
Vila-Suero, Daniel and
Limkonchotiwat, Peerat and
Marchisio, Kelly and
Leong, Wei Qi and
Susanto, Yosephine and
Ng, Raymond and
Longpre, Shayne and
Ruder, Sebastian and
Ko, Wei-Yin and
Bosselut, Antoine and
Oh, Alice and
Martins, Andre and
Choshen, Leshem and
Ippolito, Daphne and
Ferrante, Enzo and
Fadaee, Marzieh and
Ermis, Beyza and
Hooker, Sara",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.acl-long.919/",
doi = "10.18653/v1/2025.acl-long.919",
pages = "18761--18799",
ISBN = "979-8-89176-251-0"
}