2024
IDC: Boost Text-to-image Retrieval via Indirect and Direct Connections
COLING 2024main
The Dual Encoders (DE) framework maps image and text inputs into a coordinated representation space, and calculates their similarity directly. On the other hand, the Cross Attention (CA) framework performs modalities interactions after completing the feature embedding of images and text, and then ou…