2026
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
ICML 2026poster
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Although multi-turn jailbreaks are empirically effective, the structure of conversational safety failure remains insufficiently understood. In this work, we study saf…