2026
Towards Fine-grained Robustness: Attention-guided Test-time Prompt Tuning for Vision-Language Models
ICML 2026poster
Visual-Language Models (VLMs), such as CLIP, have achieved significant zero-shot performance on downstream tasks with various fine-tuning adaptation methods. However, recent studies have proven that adversarial attacks can significantly degrade the inference ability of VLMs, posing substantial risks…