2026
Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models
ICML 2026poster
Dual-encoder vision-language models (VLMs) expose a similarity interface that enables zero-shot retrieval but fails compositional constraints: queries like “umbrella and no person” retrieve images containing both, even when concept detection is reliable. We trace this to an interface-level **Bag-of-…