DiffAttack: Imperceptible and Transferable Audio Adversarial Attack via Diffusion Model
Jiayuan Chen, Yunshu Dai, Fangjun Huang
Abstract
Recently, adversarial attacks on speaker recognition systems have garnered significant interest. However, existing methods focus on injecting subtle perturbations into audio, which may compromise auditory quality. To address this problem, we propose a novel approach named DiffAttack, which employs a diffusion model for generating high-quality adversarial samples. Firstly, we extract the Mel spectrogram of the original audio. Subsequently, the Mel spectrogram is optimized to fool the speaker recognition system while preserving the high auditory quality of the attacked audio. Lastly, a conditional diffusion model is used to reconstruct the adversarial audio from the optimized Mel spectrogram. Experimental evaluations on ECAPA and ResNet, two advanced speaker recognition systems, demonstrate that our method exceeds those state-of-the-art methods in terms of attack success rate, transferability, and auditory quality.
BibTeX
@inproceedings{icassp2025_diffattackimperc,
title = {DiffAttack: Imperceptible and Transferable Audio Adversarial Attack via Diffusion Model},
author = {Jiayuan Chen and Yunshu Dai and Fangjun Huang},
booktitle = {ICASSP 2025},
year = {2025}
}