2024
Tree of Attacks: Jailbreaking Black-Box LLMs Automatically
NeurIPS 2024poster
While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed *jailbreaks*. In this work, we present *Tree of Attacks with Pruning* (TAP), an automated method for generating jailb…