← Search

Chunlong Xie

2 accepted papers

2026

Dissecting the Safety Circuit: Neuronal Intervention for Transferable Adversarial Attacks on VLMs

ICML 2026poster

The limited transferability of adversarial attacks on Vision-Language Models (VLMs) stems from their failure to navigate model-specific safety alignments, where superficial perturbations exploit surrogate-specific artifacts rather than shared safety-critical features. We reveal through linear probin…

Cited by 0SourceScholar
2025

Transstratal Adversarial Attack: Compromising Multi-Layered Defenses in Text-to-Image Models

NeurIPS 2025spotlight

Modern Text-to-Image (T2I) models deploy multi-layered defenses to block Not-Safe-For-Work (NSFW) content generation. These defenses typically include sequential layers such as prompt filters, concept erasers and image filters. While existing adversarial attacks have demonstrated vulnerabilities in…

Cited by 0SourcecodeScholar