2026
OneLIP: Unlocking and Improving Long-Text Representations of CLIP via One-Stage Adaptation
AAAI 2026technical
Contrastive Language-Image Pretraining (CLIP) has demonstrated impressive generalization on vision-language tasks by aligning images and short texts. However, its inherent 77-token length limits the capacity of capturing complex semantics in long captions. Existing long-text adaptations for CLIP typ