2023
Multi-step Jailbreaking Privacy Attacks on ChatGPT
EMNLP 2023long findings
With the rapid progress of large language models (LLMs), many downstream NLP tasks can be well solved given appropriate prompts. Though model developers and researchers work hard on dialog safety to avoid generating harmful content from LLMs, it is still challenging to steer AI-generated content (AI…