← Search

Hyrum S. Anderson

2 accepted papers

2024

Tree of Attacks: Jailbreaking Black-Box LLMs Automatically

NeurIPS 2024poster

While Large Language Models (LLMs) display versatile functionality, they continue to generate harmful, biased, and toxic content, as demonstrated by the prevalence of human-designed *jailbreaks*. In this work, we present *Tree of Attacks with Pruning* (TAP), an automated method for generating jailb…

2021

Classifying Sequences of Extreme Length with Constant Memory Applied to Malware Detection

AAAI 2021technical

Recent works within machine learning have been tackling inputs of ever increasing size, with cyber security presenting sequence classification problems of particularly extreme lengths. In the case of Windows executable malware detection, an input executable could be >=100 MB, which would translate t…