NeurIPS 2020poster92 citations

Learning Black-Box Attackers with Transferable Priors and Query Feedback

Jiancheng YANG, Yangzhou Jiang, Xiaoyang Huang, Bingbing Ni, Chenglong Zhao

Abstract

This paper addresses the challenging black-box adversarial attack problem, where only classification confidence of a victim model is available. Inspired by consistency of visual saliency between different vision models, a surrogate model is expected to improve the attack performance via transferability. By combining transferability-based and query-based black-box attack, we propose a surprisingly simple baseline approach (named SimBA++) using the surrogate model, which significantly outperforms several state-of-the-art methods. Moreover, to efficiently utilize the query feedback, we update the surrogate model in a novel learning scheme, named High-Order Gradient Approximation (HOGA). By constructing a high-order gradient computation graph, we update the surrogate model to approximate the victim model in both forward and backward pass. The SimBA++ and HOGA result in Learnable Black-Box Attack (LeBA), which surpasses previous state of the art by considerable margins: the proposed LeBA significantly reduces queries, while keeping higher attack success rates close to 100% in extensive ImageNet experiments, including attacking vision benchmarks and defensive models. Code is open source at https://github.com/TrustworthyDL/LeBA.

BibTeX
@inproceedings{NEURIPS2020_90599c8f,
 author = {YANG, Jiancheng and Jiang, Yangzhou and Huang, Xiaoyang and Ni, Bingbing and Zhao, Chenglong},
 booktitle = {Advances in Neural Information Processing Systems},
 editor = {H. Larochelle and M. Ranzato and R. Hadsell and M.F. Balcan and H. Lin},
 pages = {12288--12299},
 publisher = {Curran Associates, Inc.},
 title = {Learning Black-Box Attackers with Transferable Priors and Query Feedback},
 url = {https://proceedings.neurips.cc/paper_files/paper/2020/file/90599c8fdd2f6e7a03ad173e2f535751-Paper.pdf},
 volume = {33},
 year = {2020}
}