NeurIPS 2022accept8 citations

OccGen: Selection of Real-world Multilingual Parallel Data Balanced in Gender within Occupations

Marta R. Costa-jussà, Christine Basta, Oriol Domingo, André Niyongabo Rubungo

Abstract

This paper describes the OCCGEN toolkit, which allows extracting multilingual parallel data balanced in gender within occupations. OCCGEN can extract datasets that reflect gender diversity (beyond binary) more fairly in society to be further used to explicitly mitigate occupational gender stereotypes. We propose two use cases that extract evaluation datasets for machine translation in four high-resource languages from different linguistic families and in a low-resource African language. Our analysis of these use cases shows that translation outputs in high-resource languages tend to worsen in feminine subsets (compared to masculine). This can be explained because less attention is paid to the source sentence. Then, more attention is given to the target prefix overgeneralizing to the most frequent masculine forms.

Balanced Multilingual Data SetGenderOccupationsMachine Translation
BibTeX
@inproceedings{
costa-juss{\`a}2022occgen,
title={OccGen: Selection of Real-world Multilingual Parallel Data Balanced in Gender within Occupations},
author={Marta R. Costa-juss{\`a} and Christine Basta and Oriol Domingo and Andr{\'e} Niyongabo Rubungo},
booktitle={Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track},
year={2022},
url={https://openreview.net/forum?id=tTPVefaATp6}
}
OccGen: Selection of Real-world Multilingual Parallel Data Balanced in Gender within Occupations · NeurIPS 2022