2025
Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction
ACL 2025finding
As Large Language Models (LLMs) are widely applied in various domains, the safety of LLMs is increasingly attracting attention to avoid their powerful capabilities being misused. Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or s…