2024
Learning Fine-Grained Information Alignment for Calibrated Cross-Modal Retrieval
ICASSP 2024accepted
Masked Language Modeling (MLM) and Image-Text Matching (ITM) are always used in fusion encoder to learn the joint representation of images and text. In existing methods, the masking strategy of MLM leads to the neglect of image details during the modeling process. Meanwhile, the sampling strategy of…