← Search

Xingjun Ma*

1 accepted papers

2024

Adversarial Prompt Tuning for Vision-Language Models

ECCV 2024poster

"With the rapid advancement of multimodal learning, pre-trained Vision-Language Models (VLMs) such as CLIP have demonstrated remarkable capacities in bridging the gap between visual and language modalities. However, these models remain vulnerable to adversarial attacks, particularly in the image mod…