2024
MTA: A Lightweight Multilingual Text Alignment Model for Cross-Language Visual Word Sense Disambiguation
ICASSP 2024accepted
Visual Word Sense Disambiguation (Visual-WSD), as a sub-task of fine-grained image-text retrieval, requires a high level of language-vision understanding to capture and exploit the nuanced relationships between text and visual features. However, the cross-linguistic background only with limited cont…