2022
Robust Feature-Level Adversaries are Interpretability Tools
NeurIPS 2022accept
The literature on adversarial attacks in computer vision typically focuses on pixel-level perturbations. These tend to be very difficult to interpret. Recent work that manipulates the latent representations of image generators to create "feature-level" adversarial perturbations gives us an opportuni…