← Search

Yuyou Gan

2 accepted papers

2026

From ``Sure" to ``Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons

ICLR 2026poster

Large Vision-Language Models (LVLMs) are vulnerable to jailbreak attacks that can generate harmful content. Existing detection methods are either limited to detecting specific attack types or are too time-consuming, making them impractical for real-world deployment. To address these challenges, we p…

Cited by 0SourcecodeScholar
2025

Enhancing Adversarial Transferability with Adversarial Weight Tuning

AAAI 2025technical

Deep neural networks (DNNs) are vulnerable to adversarial examples (AEs) that mislead the model while appearing benign to human observers. A critical concern is the transferability of AEs, which enables black-box attacks without direct access to the target model. However, many previous attacks have…

Cited by 0SourcePDFScholar