2024
Nearest Neighbor Normalization Improves Multimodal Retrieval
EMNLP 2024main
Multimodal models leverage large-scale pretraining to achieve strong but still imperfect performance on tasks such as image captioning, visual question answering, and cross-modal retrieval. In this paper, we present a simple and efficient method for correcting errors in trained contrastive image-tex…