2024
PRP: Propagating Universal Perturbations to Attack Large Language Model Guard-Rails
ACL 2024long
Large language models (LLMs) are typically aligned to be harmless to humans. Unfortunately, recent work has shown that such models are susceptible to automated jailbreak attacks that induce them to generate harmful content. More recent LLMs often incorporate an additional layer of defense, a Guard M…