2024
Object-Aware Query Perturbation for Cross-Modal Image-Text Retrieval
ECCV 2024poster
"The pre-trained vision and language (V&L) models have substantially improved the performance of cross-modal image-text retrieval. In general, however, V&L models have limited retrieval performance for small objects because of the rough alignment between words and the small objects in the image. In…