NeurIPS 2017poster81 citations

Lower bounds on the robustness to adversarial perturbations

Jonathan Peck, Joris Roels, Bart Goossens, Yvan Saeys

Abstract

The input-output mappings learned by state-of-the-art neural networks are significantly discontinuous. It is possible to cause a neural network used for image recognition to misclassify its input by applying very specific, hardly perceptible perturbations to the input, called adversarial perturbations. Many hypotheses have been proposed to explain the existence of these peculiar samples as well as several methods to mitigate them. A proven explanation remains elusive, however. In this work, we take steps towards a formal characterization of adversarial perturbations by deriving lower bounds on the magnitudes of perturbations necessary to change the classification of neural networks. The bounds are experimentally verified on the MNIST and CIFAR-10 data sets.

BibTeX
@inproceedings{NIPS2017_298f95e1,
 author = {Peck, Jonathan and Roels, Joris and Goossens, Bart and Saeys, Yvan},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {I. Guyon and U. Von Luxburg and S. Bengio and H. Wallach and R. Fergus and S. Vishwanathan and R. Garnett},
 pages = {},
 publisher = {Curran Associates, Inc.},
 title = {Lower bounds on the robustness to adversarial perturbations},
 url = {https://proceedings.neurips.cc/paper_files/paper/2017/file/298f95e1bf9136124592c8d4825a06fc-Paper.pdf},
 volume = {30},
 year = {2017}
}