2023
Seeing What You Miss: Vision-Language Pre-Training With Semantic Completion Learning
CVPR 2023poster
Cross-modal alignment is essential for vision-language pre-training (VLP) models to learn the correct corresponding information across different modalities. For this purpose, inspired by the success of masked language modeling (MLM) tasks in the NLP pre-training area, numerous masked modeling tasks…