2025
Exploring Jailbreak Attacks on LLMs through Intent Concealment and Diversion
ACL 2025finding
Although large language models (LLMs) have achieved remarkable advancements, their security remains a pressing concern. One major threat is jailbreak attacks, where adversarial prompts bypass model safeguards to generate harmful or objectionable content. Researchers study jailbreak attacks to unders…