ACL 2022long35 citations

“That Is a Suspicious Reaction!”: Interpreting Logits Variation to Detect NLP Adversarial Attacks

Edoardo Mosca, Shreyash Agarwal, Javier Rando Ramírez, Georg Groh

Abstract

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extensive research in computer vision has been carried to develop reliable defense strategies. However, the same issue remains less explored in natural language processing. Our work presents a model-agnostic detector of adversarial text examples. The approach identifies patterns in the logits of the target classifier when perturbing the input text. The proposed detector improves the current state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.

BibTeX
@inproceedings{mosca-etal-2022-suspicious,
    title = "{\textquotedblleft}That Is a Suspicious Reaction!{\textquotedblright}: Interpreting Logits Variation to Detect {NLP} Adversarial Attacks",
    author = "Mosca, Edoardo  and
      Agarwal, Shreyash  and
      Rando Ram{\'i}rez, Javier  and
      Groh, Georg",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.acl-long.538/",
    doi = "10.18653/v1/2022.acl-long.538",
    pages = "7806--7816"
}
“That Is a Suspicious Reaction!”: Interpreting Logits Variation to Detect NLP Adversarial Attacks · ACL 2022