2022
Efficient Multilingual Multi-modal Pre-training through Triple Contrastive Loss
COLING 2022main
Learning visual and textual representations in the shared space from web-scale image-text pairs improves the performance of diverse vision-and-language tasks, as well as modality-specific tasks. Many attempts in this framework have been made to connect English-only texts and images, and only a few w…