← Search

Blaine Nelson

1 accepted papers

2024

Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

NeurIPS 2024poster

While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed *jailbreaks*. In this work, we present *Tree of Attacks with Pruning* (TAP), an automated method for generating jailb…