NeurIPS 2021poster22 citations

Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness

Jie Ren, Die Zhang, Yisen Wang, Lu Chen, Zhanpeng Zhou, Yiting Chen, Xu Cheng, Xin Wang

Abstract

This paper provides a unified view to explain different adversarial attacks and defense methods, i.e. the view of multi-order interactions between input variables of DNNs. Based on the multi-order interaction, we discover that adversarial attacks mainly affect high-order interactions to fool the DNN. Furthermore, we find that the robustness of adversarially trained DNNs comes from category-specific low-order interactions. Our findings provide a potential method to unify adversarial perturbations and robustness, which can explain the existing robustness-boosting methods in a principle way. Besides, our findings also make a revision of previous inaccurate understanding of the shape bias of adversarially learned features. Our code is available online at https://github.com/Jie-Ren/A-Unified-Game-Theoretic-Interpretation-of-Adversarial-Robustness.

Adversarial robustnessdeep learninggame-theoretic interaction
BibTeX
@inproceedings{
ren2021towards,
title={Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness},
author={Jie Ren and Die Zhang and Yisen Wang and Lu Chen and Zhanpeng Zhou and Yiting Chen and Xu Cheng and Xin Wang and Meng Zhou and Jie Shi and Quanshi Zhang},
booktitle={Advances in Neural Information Processing Systems},
editor={A. Beygelzimer and Y. Dauphin and P. Liang and J. Wortman Vaughan},
year={2021},
url={https://openreview.net/forum?id=fMaIxda5Y6K}
}
Towards a Unified Game-Theoretic View of Adversarial Perturbations and Robustness · NeurIPS 2021