COLING 2024main0 citations

High-Order Semantic Alignment for Unsupervised Fine-Grained Image-Text Retrieval

Rui Gao, Miaomiao Cheng, Xu Han, Wei Song

Abstract

Cross-modal retrieval is an important yet challenging task due to the semantic discrepancy between visual content and language. To measure the correlation between images and text, most existing research mainly focuses on learning global or local correspondence, failing to explore fine-grained local-global alignment. To infer more accurate similarity scores, we introduce a novel High Order Semantic Alignment (HOSA) model that can provide complementary and comprehensive semantic clues. Specifically, to jointly learn global and local alignment and emphasize local-global interaction, we employ tensor-product (t-product) operation to reconstruct one modal’s representation based on another modal’s information in a common semantic space. Such a cross-modal reconstruction strategy would significantly enhance inter-modal correlation learning in a fine-grained manner. Extensive experiments on two benchmark datasets validate that our model significantly outperforms several state-of-the-art baselines, especially in retrieving the most relevant results.

BibTeX
@inproceedings{gao-etal-2024-high,
    title = "High-Order Semantic Alignment for Unsupervised Fine-Grained Image-Text Retrieval",
    author = "Gao, Rui  and
      Cheng, Miaomiao  and
      Han, Xu  and
      Song, Wei",
    editor = "Calzolari, Nicoletta  and
      Kan, Min-Yen  and
      Hoste, Veronique  and
      Lenci, Alessandro  and
      Sakti, Sakriani  and
      Xue, Nianwen",
    booktitle = "Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)",
    month = may,
    year = "2024",
    address = "Torino, Italia",
    publisher = "ELRA and ICCL",
    url = "https://aclanthology.org/2024.lrec-main.714/",
    pages = "8155--8165"
}
High-Order Semantic Alignment for Unsupervised Fine-Grained Image-Text Retrieval · COLING 2024