← Search

Darina Koishigarina

2 accepted papers

2026

CLIP Behaves like a Bag-of-Words Model Cross-modally but not Uni-modally

ICLR 2026poster

CLIP (Contrastive Language-Image Pretraining) has become a popular choice for various downstream tasks. However, recent studies have questioned its ability to represent compositional concepts effectively. These works suggest that CLIP often acts like a bag-of-words (BoW) model, interpreting images a…

Cited by 0SourcecodeScholar