2024
Long-CLIP: Unlocking the Long-Text Capability of CLIP
ECCV 2024poster
"Contrastive Language-Image Pre-training (CLIP) has been the cornerstone for zero-shot classification, text-image retrieval, and text-image generation by aligning image and text modalities. Despite its widespread adoption, a significant limitation of CLIP lies in the inadequate length of text input.…