AAAI 2025technical0 citations

Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy

Jian-Ping Mei, Weibin Zhang, Jie Chen, Xuyun Zhang, Tiantian Zhu

Abstract

Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings.

BibTeX
@article{Mei_Zhang_Chen_Zhang_Zhu_2025, title={Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy}, volume={39}, url={https://ojs.aaai.org/index.php/AAAI/article/view/32041}, DOI={10.1609/aaai.v39i1.32041}, abstractNote={Malicious users attempt to replicate commercial models functionally at low cost by training a clone model with query responses. It is challenging to timely prevent such model-stealing attacks to achieve strong protection and maintain utility. In this paper, we propose a novel non-parametric detector called Account-aware Distribution Discrepancy (ADD) to recognize queries from malicious users by leveraging account-wise local dependency. We formulate each class as a Multivariate Normal distribution (MVN) in the feature space and measure the malicious score as the sum of weighted class-wise distribution discrepancy. The ADD detector is combined with random-based prediction poisoning to yield a plug-and-play defense module named D-ADD for image classification models. Results of extensive experimental studies show that D-ADD achieves strong defense against different types of attacks with little interference in serving benign users for both soft and hard-label settings.}, number={1}, journal={Proceedings of the AAAI Conference on Artificial Intelligence}, author={Mei, Jian-Ping and Zhang, Weibin and Chen, Jie and Zhang, Xuyun and Zhu, Tiantian}, year={2025}, month={Apr.}, pages={604-611} }
Defense Against Model Stealing Based on Account-Aware Distribution Discrepancy · AAAI 2025