2024
XVD: Cross-Vocabulary Differentiable Training for Generative Adversarial Attacks
COLING 2024main
An adversarial attack to a text classifier consists of an input that induces the classifier into an incorrect class prediction, while retaining all the linguistic properties of correctly-classified examples. A popular class of adversarial attacks exploits the gradients of the victim classifier to tr…