← Search

Yifeng Feng

1 accepted papers

2025

Rewrite to Jailbreak: Discover Learnable and Transferable Implicit Harmfulness Instruction

ACL 2025finding

As Large Language Models (LLMs) are widely applied in various domains, the safety of LLMs is increasingly attracting attention to avoid their powerful capabilities being misused. Existing jailbreak methods create a forced instruction-following scenario, or search adversarial prompts with prefix or s…