2025
Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
NeurIPS 2025poster
Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. While existing methods typically achieve targeted attacks by aligning global features—such as CLIP’s [CLS] token—between adversarial and target samples, they often overlook the rich local information enc…