← Search

Suhuang Wu

1 accepted papers

2025

Monte Carlo Tree Search Based Prompt Autogeneration for Jailbreak Attacks against LLMs

COLING 2025main

Jailbreak attacks craft specific prompts or append adversarial suffixes to prompts, thereby inducing language models to generate harmful or unethical content and bypassing the model’s safety guardrails. With the recent blossom of large language models (LLMs), there’s a growing focus on jailbreak att…