Adversarial Knowledge Transfer for Black-Box Model Inversion Attack
Xinhao Liu, Zetao Lin, Yingzhao Jiang, Qiao Yan
Abstract
Recent advancements in model inversion attacks have raised privacy concerns, exploiting access to models to reconstruct private training data from inputs and outputs. These attacks are categorized as white-box, black-box, or label-only based on access level. We propose a new black-box model inversion attack, Label-Controlled Adversarial Knowledge Transfer (L-AdKT). L-AdKT leverages adversarial training with a Generative Adversarial Network (GAN) and a substitute model to extract knowledge from the target model. The substitute model minimizes discrepancies with the target model while guiding the generator to produce realistic samples. This approach enables white-box techniques to be applied in black-box settings. Experiments show that L-AdKT outperforms state-of-the-art black-box attacks by over 20% across benchmarks and remains robust against various defense mechanisms.
BibTeX
@inproceedings{icassp2025_adversarialknowl,
title = {Adversarial Knowledge Transfer for Black-Box Model Inversion Attack},
author = {Xinhao Liu and Zetao Lin and Yingzhao Jiang and Qiao Yan},
booktitle = {ICASSP 2025},
year = {2025}
}