2026
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
ICLR 2026poster
Vision–language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to image-based jailbreaks crafted to induce harmful responses. E…