← Search

Danil Malaev

1 accepted papers

2023

Layerwise universal adversarial attack on NLP models

ACL 2023findings

In this work, we examine the vulnerability of language models to universal adversarial triggers (UATs). We propose a new white-box approach to the construction of layerwise UATs (LUATs), which searches the triggers by perturbing hidden layers of a network. On the example of three transformer models…