2025
ToxicTextCLIP: Text-Based Poisoning and Backdoor Attacks on CLIP Pre-training
NeurIPS 2025poster
The Contrastive Language-Image Pretraining (CLIP) model has significantly advanced vision-language modeling by aligning image-text pairs from large-scale web data through self-supervised contrastive learning. Yet, its reliance on uncurated Internet-sourced data exposes it to data poisoning and backd…