← Search

Joe D. Menke

1 accepted papers

2024

Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters

NeurIPS 2024poster

Large Language Models (LLMs) are typically harmless but remain vulnerable to carefully crafted prompts known as ``jailbreaks'', which can bypass protective measures and induce harmful behavior. Recent advancements in LLMs have incorporated moderation guardrails that can filter outputs, which trigger…

Cited by 14SourcePDFScholar