2024
The Hard Positive Truth about Vision-Language Compositionality
ECCV 2024poster
"Several benchmarks have concluded that our best vision-language models (, CLIP) are lacking in compositionality. Given an image, these benchmarks probe a model’s ability to identify its associated caption amongst a set of compositional distractors. In response, a surge of recent proposals show impr…