2024
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
ICLR 2024spotlight
While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to *prompt injection attacks*: malicious third party prompts that subvert the intent of the system designer. To help researchers study this problem, we present a dataset of over 563,000…