2024
Integration of Global and Local Representations for Fine-grained Cross-modal Alignment
ECCV 2024poster
"Fashion is one of the representative domains of fine-grained Vision-Language Pre-training (VLP) involving a large number of images and text. Previous fashion VLP research has proposed various pre-training tasks to account for fine-grained details in multimodal fusion. However, fashion VLP research…