← Search

Li jiaye

2 accepted papers

2025

Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation

CVPR 2025poster

Contrastive language-image pretraining models such as CLIP have demonstrated remarkable performance in various text-image alignment tasks. However, the inherent 77-token input limitation and reliance on predominantly short-text training data restrict its ability to handle long-text tasks effectively…

Cited by 0SourcePDFScholar