Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models
Recent vision-language models (VLMs), such as CLIP, have demonstrated remarkable transferability across a wide range of downstream tasks by effectively leveraging the joint text-image embedding space, even with only a few data samples. Despite their impressive performance, these models remain vulner…