2023
Fine-Grained Image-Text Matching by Cross-Modal Hard Aligning Network
CVPR 2023poster
Current state-of-the-art image-text matching methods implicitly align the visual-semantic fragments, like regions in images and words in sentences, and adopt cross-attention mechanism to discover fine-grained cross-modal semantic correspondence. However, the cross-attention mechanism may bring redun…