2022
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality
CVPR 2022poster
We present a novel task and dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning, which we call Winoground. Given two images and two captions, the goal is to match them correctly--but crucially, both captions contain a completely identi…