ICASSP 2025accepted0 citations

Bridging the Modality Gap for Speech-image Retrieval with Text Supervision

Yuting Yang, Lifeng Zhou, Yuke Li, Guodong Ma

Abstract

In recent years, while the performance of speech-image retrieval has improved significantly, it still lags behind that of image-text retrieval. Leveraging the text modality to enhance speech-image retrieval remains a promising research direction. In this paper, we propose to leverage text supervision to facilitate the alignment between speech and image feature spaces via an automatic speech recognition (ASR) auxiliary task. Specifically, our model is trained with a multi-task learning framework, which combines ASR and speech-image contrastive learning tasks. On this basis, benefiting from the ASR module to obtain text modality information, we introduce an ensemble mechanism to further enhance speech-image retrieval performance. Experimental results on two publicly available datasets Flickr8k and SpokenCOCO indicate that the introduction of the ASR task improves the mean R@1 by 2.7% and 2%, respectively, compared to the previous state-of-the-art (SOTA) method. Furthermore, the results demonstrate that the ensemble mechanism further enhances the performance of speech-image retrieval.

BibTeX
@inproceedings{icassp2025_bridgingthemodal,
  title = {Bridging the Modality Gap for Speech-image Retrieval with Text Supervision},
  author = {Yuting Yang and Lifeng Zhou and Yuke Li and Guodong Ma},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Bridging the Modality Gap for Speech-image Retrieval with Text Supervision · ICASSP 2025