← Search

Gaozheng Pei

4 accepted papers

2026

Localize and Neutralize: Gradient-Guided Token Suppression Against Visual Prompt Injection Attack

ICML 2026poster

Adversarial images pose a severe security threat to multimodal large language models through prompt injection. Existing defenses largely lack a principled understanding of the underlying mechanisms and struggle to balance efficiency and fidelity. In this work, we show that successful adversarial att…

Cited by 0SourceScholar
2025

Diffusion-based Adversarial Purification from the Perspective of the Frequency Domain

ICML 2025spotlight

The diffusion-based adversarial purification methods attempt to drown adversarial perturbations into a part of isotropic noise through the forward process, and then recover the clean images through the reverse process. Due to the lack of distribution information about adversarial perturbations in th…

Cited by 0SourcePDFScholar
2025

Divide and Conquer: Heterogeneous Noise Integration for Diffusion-based Adversarial Purification

CVPR 2025poster

Existing diffusion-based purification methods aim to disrupt adversarial perturbations by introducing a certain amount of noise through a forward diffusion process, followed by a reverse process to recover clean examples. However, this approach is fundamentally flawed: the uniform operation of the f…

Cited by 2SourcePDFScholar
2025

Exploring Query Efficient Data Generation Towards Data-Free Model Stealing in Hard Label Setting

AAAI 2025technical

Data-free model stealing involves replicating the functionality of a target model into a substitute model without accessing the target model's structure, parameters, or training data. Instead, the adversary can only access the target model's predictions for generated samples. Once the substitute mod…

Cited by 1SourcePDFScholar