← Search

Ema Borevković

1 accepted papers

2024

Query-Based Adversarial Prompt Generation

NeurIPS 2024poster

Recent work has shown it is possible to construct adversarial examples that cause aligned language models to emit harmful strings or perform harmful behavior. Existing attacks work either in the white-box setting (with full access to the model weights), or through _transferability_: the phenomenon t…

Cited by 32SourcePDFScholar