← Search

Aldo Zaimi

2 accepted papers

2026

CLIP Is Shortsighted: Paying Attention Beyond the First Sentence

CVPR 2026

CLIP models learn transferable multi-modal features via image-text contrastive learning on internet-scale data. They are widely used in zero-shot classification, multi-modal retrieval, text-to-image diffusion, and as image encoders in large vision-language models. However, CLIP's pretraining is domi

Cited by 0SourcecodeScholar
2024

CableInspect-AD: An Expert-Annotated Anomaly Detection Dataset

NeurIPS 2024poster

Machine learning models are increasingly being deployed in real-world contexts. However, systematic studies on their transferability to specific and critical applications are underrepresented in the research literature. An important example is visual anomaly detection (VAD) for robotic power line in…