2025
Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs
COLING 2025main
Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. With the recent blossom of large language models (LLMs), there’s a growing focus on jailbreak att…