ICLR 2023poster17 citations

Why adversarial training can hurt robust accuracy

Jacob Clarysse, Julia Hörrmann, Fanny Yang

Abstract

Machine learning classifiers with high test accuracy often perform poorly under adversarial attacks. It is commonly believed that adversarial training alleviates this issue. In this paper, we demonstrate that, surprisingly, the opposite can be true for a natural class of perceptible perturbations --- even though adversarial training helps when enough data is available, it may in fact hurt robust generalization in the small sample size regime. We first prove this phenomenon for a high-dimensional linear classification setting with noiseless observations. Using intuitive insights from the proof, we could surprisingly find perturbations on standard image datasets for which this behavior persists. Specifically, it occurs for perceptible attacks that effectively reduce class information such as object occlusions or corruptions.

Adversarial trainingLearning TheoryRobust generalisation
BibTeX
@inproceedings{
clarysse2023why,
title={Why adversarial training can hurt robust accuracy},
author={Jacob Clarysse and Julia H{\"o}rrmann and Fanny Yang},
booktitle={The Eleventh International Conference on Learning Representations },
year={2023},
url={https://openreview.net/forum?id=-CA8yFkPc7O}
}
Why adversarial training can hurt robust accuracy · ICLR 2023