EMNLP 2024industry8 citations

Prompt Leakage effect and mitigation strategies for multi-turn LLM Applications

Divyansh Agarwal, Alexander Fabbri, Ben Risher, Philippe Laban, Shafiq Joty, Chien-Sheng Wu

Abstract

Prompt leakage poses a compelling security and privacy threat in LLM applications. Leakage of system prompts may compromise intellectual property, and act as adversarial reconnaissance for an attacker. A systematic evaluation of prompt leakage threats and mitigation strategies is lacking, especially for multi-turn LLM interactions. In this paper, we systematically investigate LLM vulnerabilities against prompt leakage for 10 closed- and open-source LLMs, across four domains. We design a unique threat model which leverages the LLM sycophancy effect and elevates the average attack success rate (ASR) from 17.7% to 86.2% in a multi-turn setting. Our standardized setup further allows dissecting leakage of specific prompt contents such as task instructions and knowledge documents. We measure the mitigation effect of 7 black-box defense strategies, along with finetuning an open-source model to defend against leakage attempts. We present different combination of defenses against our threat model, including a cost analysis. Our study highlights key takeaways for building secure LLM applications and provides directions for research in multi-turn LLM interactions.

BibTeX
@inproceedings{agarwal-etal-2024-prompt,
    title = "Prompt Leakage effect and mitigation strategies for multi-turn {LLM} Applications",
    author = "Agarwal, Divyansh  and
      Fabbri, Alexander  and
      Risher, Ben  and
      Laban, Philippe  and
      Joty, Shafiq  and
      Wu, Chien-Sheng",
    editor = "Dernoncourt, Franck  and
      Preo{\c{t}}iuc-Pietro, Daniel  and
      Shimorina, Anastasia",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track",
    month = nov,
    year = "2024",
    address = "Miami, Florida, US",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-industry.94/",
    doi = "10.18653/v1/2024.emnlp-industry.94",
    pages = "1255--1275"
}
Prompt Leakage effect and mitigation strategies for multi-turn LLM Applications · EMNLP 2024