← Search

Shengli Song*

1 accepted papers

2024

Improving Vision and Language Concepts Understanding with Multimodal Counterfactual Samples

ECCV 2024poster

"Vision and Language (VL) models have achieved remarkable performance in a variety of multimodal learning tasks. The success of these models is attributed to learning a joint and aligned representation space of visual and text. However, recent popular VL models still struggle with concepts understan…