← Search

Sergey Berezin

2 accepted papers

2025

The TIP of the Iceberg: Revealing a Hidden Class of Task-in-Prompt Adversarial Attacks on LLMs

ACL 2025long

We present a novel class of jailbreak adversarial attacks on LLMs, termed Task-in-Prompt (TIP) attacks. Our approach embeds sequence-to-sequence tasks (e.g., cipher decoding, riddles, code execution) into the model’s prompt to indirectly generate prohibited inputs. To systematically assess the effec…

Cited by 0SourcePDFScholar
2023

No offence, Bert - I insult only humans! Multilingual sentence-level attack on toxicity detection networks

EMNLP 2023short findings

We introduce a simple yet efficient sentence-level attack on black-box toxicity detector models. By adding several positive words or sentences to the end of a hateful message, we are able to change the prediction of a neural network and pass the toxicity detection system check. This approach is show…

Cited by 0SourceScholar