NeurIPS 2022accept30 citations
Gradient Methods Provably Converge to Non-Robust Networks
Gal Vardi, Gilad Yehudai, Ohad Shamir
Abstract
Despite a great deal of research, it is still unclear why neural networks are so susceptible to adversarial examples. In this work, we identify natural settings where depth-$2$ ReLU networks trained with gradient flow are provably non-robust (susceptible to small adversarial $\ell_2$-perturbations), even when robust networks that classify the training dataset correctly exist. Perhaps surprisingly, we show that the well-known implicit bias towards margin maximization induces bias towards non-robust networks, by proving that every network which satisfies the KKT conditions of the max-margin problem is non-robust.
implicit biasdeep learning theoryrobustness
BibTeX
@inproceedings{
vardi2022gradient,
title={Gradient Methods Provably Converge to Non-Robust Networks},
author={Gal Vardi and Gilad Yehudai and Ohad Shamir},
booktitle={Advances in Neural Information Processing Systems},
editor={Alice H. Oh and Alekh Agarwal and Danielle Belgrave and Kyunghyun Cho},
year={2022},
url={https://openreview.net/forum?id=XDZhagjfMP}
}