2025
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
CVPR 2025highlight
Contrastive Language-Image Pre-training (CLIP) \citep radford2021learning has emerged as a pivotal model in computer vision and multimodal learning, achieving state-of-the-art performance at aligning visual and textual representations through contrastive learning. However, CLIP struggles with poten…