2025
Semantic and Expressive Variations in Image Captions Across Languages
CVPR 2025poster
Most vision-language models today are primarily trained on English image-text pairs, with non-English pairs often filtered out. Evidence from cross-cultural psychology suggests that this approach will bias models against perceptual modes exhibited by people who speak other (non-English) languages. W…