ICASSP 2025accepted0 citations

Unveiling Local Well-posedness Influence for Cross-modal Person Re-Identification

Yumeng Yang, Guannan Dong, Aichun Zhu, Mingcheng Ni, Yifeng Li

Abstract

The existing cross-modal retrieval methods trend toward the conventional multi-modal alignment while ignoring the localization bias caused by visual hallucination, including color pollution and appearance-like occlusion due to uncontrollable factors such as weather, illumination, and occlusion. This feature blinding misleads the model to lock in the pseudo-real position and further leads to local unmatched. To this end, we discuss cross-modal local alignment well-posedness by making a phased local modal-masking to calibrate the undisturbed actual local alignment from entity, attribute, and appearance. Specifically, we introduce a mask-based local well-posedness modeling (MLWM) strategy, including text-based entity masking (TEM), text-based attribute-specific masking (TAM), and image-based appearance masking (IAM) to phased collaboratively consider image prompting-based text entities, image prompting-based text attributes, and text prompting-based appearance inference contrast, respectively. Finally, we dynamically optimize the weights of positively correlated image-text pairs by comparing the similarity between original and reconstructed features. Experimental results demonstrate that our method is effective on three public datasets.

BibTeX
@inproceedings{icassp2025_unveilinglocalwe,
  title = {Unveiling Local Well-posedness Influence for Cross-modal Person Re-Identification},
  author = {Yumeng Yang and Guannan Dong and Aichun Zhu and Mingcheng Ni and Yifeng Li},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Unveiling Local Well-posedness Influence for Cross-modal Person Re-Identification · ICASSP 2025