2026
Structured Multi-step Jailbreaking under a Hamiltonian Generative Formulation
ICML 2026poster
Recent work shows that even safety aligned large language models (LLM) can be pushed into unsafe behavior by carefully crafted jailbreak prompts. Existing jailbreaking attack methods often rely on disfluent or incoherent prompts, which limit their success and make them easy to detect. We introduce S…