← Search

Gengshen Zhang

1 accepted papers

2025

FG-CLIP: Fine-Grained Visual and Textual Alignment

ICML 2025poster

Contrastive Language-Image Pre-training (CLIP) excels in multimodal tasks such as image-text retrieval and zero-shot classification but struggles with fine-grained understanding due to its focus on coarse-grained short captions. To address this, we propose Fine-Grained CLIP (FG-CLIP), which enhances…