Unraveling the Mystery: Defending Against Jailbreak Attacks Via Unearthing Real Intention
As Large Language Models (LLMs) become more advanced, the security risks they pose also increase. Ensuring that LLM behavior aligns with human values, particularly in mitigating jailbreak attacks with elusive and implicit intentions, has become a significant challenge. To address this issue, we prop…