2024
GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation
EMNLP 2024main
Research on jailbreaking has been valuable for testing and understanding the safety and security issues of large language models (LLMs). In this paper, we introduce Iterative Refinement Induced Self-Jailbreak (IRIS), a novel approach that leverages the reflective capabilities of LLMs for jailbreakin…