SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models
Margaret Mitchell, Giuseppe Attanasio, Ioana Baldini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton
Abstract
Large Language Models (LLMs) reproduce and exacerbate the social biases present in their training data, and resources to quantify this issue are limited. While research has attempted to identify and mitigate such biases, most efforts have been concentrated around English, lagging the rapid advancement of LLMs in multilingual settings. In this paper, we introduce a new multilingual parallel dataset SHADES to help address this issue, designed for examining culturally-specific stereotypes that may be learned by LLMs. The dataset includes stereotypes from 20 regions around the world and 16 languages, spanning multiple identity categories subject to discrimination worldwide. We demonstrate its utility in a series of exploratory evaluations for both “base” and “instruction-tuned” language models. Our results suggest that stereotypes are consistently reflected across models and languages, with some languages and models indicating much stronger stereotype biases than others.
BibTeX
@inproceedings{mitchell-etal-2025-shades,
title = "{SHADES}: Towards a Multilingual Assessment of Stereotypes in Large Language Models",
author = "Mitchell, Margaret and
Attanasio, Giuseppe and
Baldini, Ioana and
Clinciu, Miruna and
Clive, Jordan and
Delobelle, Pieter and
Dey, Manan and
Hamilton, Sil and
Dill, Timm and
Doughman, Jad and
Dutt, Ritam and
Ghosh, Avijit and
Forde, Jessica Zosa and
Holtermann, Carolin and
Kaffee, Lucie-Aim{\'e}e and
Laud, Tanmay and
Lauscher, Anne and
Lopez-Davila, Roberto L and
Masoud, Maraim and
Nangia, Nikita and
Ovalle, Anaelia and
Pistilli, Giada and
Radev, Dragomir and
Savoldi, Beatrice and
Raheja, Vipul and
Qin, Jeremy and
Ploeger, Esther and
Subramonian, Arjun and
Dhole, Kaustubh and
Sun, Kaiser and
Djanibekov, Amirbek and
Mansurov, Jonibek and
Yin, Kayo and
Cueva, Emilio Villa and
Mukherjee, Sagnik and
Huang, Jerry and
Shen, Xudong and
Gala, Jay and
Al-Ali, Hamdan and
Tair Djanibekov and
Mukhituly, Nurdaulet and
Nie, Shangrui and
Sharma, Shanya and
Stanczak, Karolina and
Szczechla, Eliza and
Timponi Torrent, Tiago and
Tunuguntla, Deepak and
Viridiano, Marcelo and
Van Der Wal, Oskar and
Yakefu, Adina and
N{\'e}v{\'e}ol, Aur{\'e}lie and
Zhang, Mike and
Zink, Sydney and
Talat, Zeerak",
editor = "Chiruzzo, Luis and
Ritter, Alan and
Wang, Lu",
booktitle = "Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
month = apr,
year = "2025",
address = "Albuquerque, New Mexico",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.naacl-long.600/",
pages = "11995--12041",
ISBN = "979-8-89176-189-6"
}