2025
Feint and Attack: Jailbreaking and Protecting LLMs via Attention Distribution Modeling
IJCAI 2025
Most jailbreak methods for large language models (LLMs) focus on superficially improving attack success through manually defined rules. However, they fail to uncover the underlying mechanisms within target LLMs that explain why an attack succeeds or fails. In this paper, we propose investigating the