2026
Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMs
ICML 2026poster
The limited transferability of adversarial attacks on Vision-Language Models (VLMs) stems from their failure to navigate model-specific safety alignments, where superficial perturbations exploit surrogate-specific artifacts rather than shared safety-critical features. We reveal through linear probin…