ACL 2022findings34 citations

Improving the Adversarial Robustness of NLP Models by Information Bottleneck

Cenyuan Zhang, Xiang Zhou, Yixin Wan, Xiaoqing Zheng, Kai-Wei Chang, Cho-Jui Hsieh

Abstract

Existing studies have demonstrated that adversarial examples can be directly attributed to the presence of non-robust features, which are highly predictive, but can be easily manipulated by adversaries to fool NLP models. In this study, we explore the feasibility of capturing task-specific robust features, while eliminating the non-robust ones by using the information bottleneck theory. Through extensive experiments, we show that the models trained with our information bottleneck-based method are able to achieve a significant improvement in robust accuracy, exceeding performances of all the previously reported defense methods while suffering almost no performance drop in clean accuracy on SST-2, AGNEWS and IMDB datasets.

BibTeX
@inproceedings{zhang-etal-2022-improving,
    title = "Improving the Adversarial Robustness of {NLP} Models by Information Bottleneck",
    author = "Zhang, Cenyuan  and
      Zhou, Xiang  and
      Wan, Yixin  and
      Zheng, Xiaoqing  and
      Chang, Kai-Wei  and
      Hsieh, Cho-Jui",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2022",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.findings-acl.284/",
    doi = "10.18653/v1/2022.findings-acl.284",
    pages = "3588--3598"
}
Improving the Adversarial Robustness of NLP Models by Information Bottleneck · ACL 2022