2025
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
ICLR 2025poster
The integration of new modalities into frontier AI systems offers exciting capabilities, but also increases the possibility such systems can be adversarially manipulated in undesirable ways. In this work, we focus on a popular class of vision-language models (VLMs) that generate text outputs conditi…