← Search

Yunan Zeng

2 accepted papers

2024

Investigating Compositional Challenges in Vision-Language Models for Visual Grounding

CVPR 2024highlight

Pre-trained vision-language models (VLMs) have achieved high performance on various downstream tasks which have been widely used for visual grounding tasks in a weakly supervised manner. However despite the performance gains contributed by large vision and language pre-training we find that state-of…

2022

MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text Matching

NeurIPS 2022accept

Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which are trained on millions or billions of paired images and texts. Different from them, this paper studies a new scenario as unpaired image-text matching, in which paired images and tex…

Cited by 26SourcePDFScholar