2020
Proper Network Interpretability Helps Adversarial Robustness in Classification
ICML 2020poster
Recent works have empirically shown that there exist adversarial examples that can be hidden from neural network interpretability (namely, making network interpretation maps visually similar), or interpretability is itself susceptible to adversarial attacks. In this paper, we theoretically show that…