2025
Mitigate the Gap: Improving Cross-Modal Alignment in CLIP
ICLR 2025poster
Contrastive Language--Image Pre-training (CLIP) has manifested remarkable improvements in zero-shot classification and cross-modal vision-language tasks. Yet, from a geometrical point of view, the CLIP embedding space has been found to have a pronounced modality gap. This gap renders the embedding s…