2024
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
COLING 2024main
Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited…