2025
Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language Models
NeurIPS 2025poster
Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt…