← Search

Yoyo Minzhi Zhang

1 accepted papers

2023

When are Lemons Purple? The Concept Association Bias of Vision-Language Models

EMNLP 2023long main

Large-scale vision-language models such as CLIP have shown impressive performance on zero-shot image classification and image-to-text retrieval. However, such performance does not realize in tasks that require a finer-grained correspondence between vision and language, such as Visual Question Answer…

Cited by 0SourceScholar