2024
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
NeurIPS 2024poster
Large Language Models (LLMs) are typically harmless but remain vulnerable to carefully crafted prompts known as ``jailbreaks'', which can bypass protective measures and induce harmful behavior. Recent advancements in LLMs have incorporated moderation guardrails that can filter outputs, which trigger…