← Search

Atharva Roshan Naik

1 accepted papers

2024

Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks

COLING 2024main

Recent explorations with commercial Large Language Models (LLMs) have shown that non-expert users can jailbreak LLMs by simply manipulating their prompts; resulting in degenerate output behavior, privacy and security breaches, offensive outputs, and violations of content regulator policies. Limited…