Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial Attacks
Pre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of VLMs become increasingly widespread, their potential safety and robustness issues raise conc…