Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
We discover a novel and surprising phenomenon of unintentional misalignment in reasoning language models (RLMs), which we call \textit{self-jailbreaking}. Specifically, after benign reasoning training on math or code domains, RLMs will use multiple strategies to circumvent their own safety guardrail…