Improving Vision and Language Concepts Understanding with Multimodal Counterfactual Samples
"Vision and Language (VL) models have achieved remarkable performance in a variety of multimodal learning tasks. The success of these models is attributed to learning a joint and aligned representation space of visual and text. However, recent popular VL models still struggle with concepts understan…