2025
Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?
ACL 2025finding
Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the questi…