2022
Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
EMNLP 2022main
Research in massively multilingual image captioning has been severely hampered by a lack of high-quality evaluation datasets. In this paper we present the Crossmodal-3600 dataset (XM3600 in short), a geographically diverse set of 3600 images annotated with human-generated reference captions in 36 la…