Leveraging Large Vision-Language Model as User Intent-Aware Encoder for Composed Image Retrieval
Composed Image Retrieval (CIR) aims to retrieve target images from candidate set using a hybrid-modality query consisting of a reference image and a relative caption that describes the user intent. Recent studies attempt to utilize Vision-Language Pre-training Models (VLPMs) with various fusion stra…