ICASSP 2024accepted0 citations

MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval

Tianle Lv, Shuang Li, Jiaxu Leng, Xinbo Gao

Abstract

Text-to-image person retrieval aims to recognize target pedestrians based on specified text. Existing methods mainly obtain image and text features separately through distinct feature extractors, subsequently embedding them into a unified feature space and calculating their similarity. Despite great success, current methods still suffer from the lack of information interaction between images and text. To address this issue, we propose Mutual-guidance Representation Learning (MGRL) for text-to-image person retrieval, which captures the key features for matching via text-image information interaction. Accordingly, our MGRL consists of two customized modules: iterative text-guided feature extraction (ITFE) and vision-assisted specific mask complement (VSMC). Specifically, ITFE is first designed to extract the matching information between the text and the image concerning the local feature attention of the target pedestrians by iterative text guidance. Then, to further ensure the image features extracted by ITFE contain the text description, VSMC is designed to utilize the extracted image features to help complete masked text where the mask is difficult to complete with only unmasked text information. Experiments are conducted on CUHK-PEDES and ICFG-PEDES datasets, and experimental results demonstrate the superiority of the proposed MGRL.

BibTeX
@inproceedings{icassp2024_mgrlmutualguidan,
  title = {MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval},
  author = {Tianle Lv and Shuang Li and Jiaxu Leng and Xinbo Gao},
  booktitle = {ICASSP 2024},
  year = {2024}
}
MGRL: Mutual-Guidance Representation Learning for Text-to-Image Person Retrieval · ICASSP 2024