← Search

Shreyas Chandrashekaran

1 accepted papers

2024

PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails

ACL 2024long

Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks that induce them to generate harmful content. More recent LLMs often incorporate an additional layer of defense, a Guard M…