2024
SILC: Improving Vision Language Pretraining with Self-Distillation
ECCV 2024poster
"Image-Text pretraining on web-scale image caption datasets has become the default recipe for open vocabulary classification and retrieval models thanks to the success of CLIP and its variants. Several works have also used CLIP features for dense prediction tasks and have shown the emergence of open…