2025
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
ACL 2025long
Current studies have exposed the risk of Large Language Models (LLMs) generating harmful content by jailbreak attacks. However, they overlook that the direct generation of harmful content from scratch is more difficult than inducing LLM to calibrate benign content into harmful forms.In our study, we…