2024
ComCLIP: Training-Free Compositional Image and Text Matching
NAACL 2024long
Contrastive Language-Image Pretraining (CLIP) has demonstrated great zero-shot performance for matching images and text. However, it is still challenging to adapt vision-language pretrained models like CLIP to compositional image and text matching — a more challenging image and text matching task re…