2025
FineLIP: Extending CLIP's Reach via Fine-Grained Alignment with Longer Text Inputs
CVPR 2025poster
As a pioneering vision-language model, CLIP (Contrastive Language-Image Pre-training) has achieved significant success across various domains and a wide range of downstream vision-language tasks. However, the text encoders in popular CLIP models are limited to processing only 77 text tokens, which c…