2021
COOKIE: Contrastive Cross-Modal Knowledge Sharing Pre-Training for Vision-Language Representation
ICCV 2021poster
There has been a recent surge of interest in cross-modal pre-training. However, existed approaches pre-train a one-stream model to learn joint vision-language representation, which suffers from calculation explosion when conducting cross-modal retrieval. In this work, we propose the Contrastive Cros…