2022
Vision-Language Pre-Training With Triple Contrastive Learning
CVPR 2022poster
Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attributed to its capability in maximizing the mutual information (MI) between an image and its matched text. However, simply…