2022
Conditioned Masked Language and Image Modeling for Image-Text Dense Retrieval
EMNLP 2022finding
Image-text retrieval is a fundamental cross-modal task that takes image/text as a query to retrieve relevant data of another type. The large-scale two-stream pre-trained models like CLIP have achieved tremendous success in this area. They embed the images and texts into instance representations with…