2025
Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLM
AAAI 2025technical
Despite being empowered with alignment mechanisms, large language models (LLMs) are increasingly vulnerable to emerging jailbreak attacks that can compromise their alignment mechanisms. This vulnerability poses significant risks to real-world applications. Existing work faces challenges in both tra…