2022
Disentangling Visual and Written Concepts in CLIP
CVPR 2022oral
The CLIP network measures the similarity between natural text and images; in this work, we investigate the entanglement of the representation of word images and natural images in its image encoder. First, we find that the image encoder has an ability to match word images with natural images of scene…