ICASSP 2023accepted0 citations

Semantic-Preserving Augmentation for Robust Image-Text Retrieval

Sunwoo Kim, Kyuhong Shim, Luong Trung Nguyen, Byonghyo Shim

Abstract

Image-text retrieval is a task to search for the proper textual descriptions of the visual world and vice versa. One challenge of this task is the vulnerability to input image/text corruptions. Such corruptions are often unobserved during the training, and degrade the retrieval model’s decision quality substantially. In this paper, we propose a novel image-text retrieval technique, referred to as robust visual semantic embedding (RVSE), which consists of novel image-based and text-based augmentation techniques called semantic-preserving augmentation for image (SPAug-I) and text (SPAug-T). Since SPAug-I and SPAug-T change the original data in a way that its semantic information is preserved, we enforce the feature extractors to generate semantic-aware embedding vectors regardless of the corruption, improving the model’s robustness significantly. From extensive experiments using benchmark datasets, we show that RVSE outperforms conventional retrieval schemes in terms of image-text retrieval performance.

BibTeX
@inproceedings{icassp2023_semanticpreservi,
  title = {Semantic-Preserving Augmentation for Robust Image-Text Retrieval},
  author = {Sunwoo Kim and Kyuhong Shim and Luong Trung Nguyen and Byonghyo Shim},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Semantic-Preserving Augmentation for Robust Image-Text Retrieval · ICASSP 2023