2022
ViSTA: Vision and Scene Text Aggregation for Cross-Modal Retrieval
CVPR 2022poster
Visual appearance is considered to be the most important cue to understand images for cross-modal retrieval, while sometimes the scene text appearing in images can provide valuable information to understand the visual semantics. Most of existing cross-modal retrieval approaches ignore the usage of s…