2026
JULI: Jailbreak Large Language Models by Self-Introspection
ICLR 2026poster
Large Language Models (LLMs) are trained with safety alignment to prevent generating malicious content. Although some attacks have highlighted vulnerabilities in these safety-aligned LLMs, they typically have limitations, such as necessitating access to the model weights or the generation process. S…