2025
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
ICCV 2025poster
Driven by large-scale contrastive vision-language pre-trained models such as CLIP, recent advancements in the image-text matching task have achieved remarkable success in representation learning. Due to image-level visual-language alignment, CLIP falls short in understanding fine-grained details suc…