ICASSP 2025accepted0 citations

Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring

Yujun Zhang, Yanqu Chen, Jiakai Wang, Jin Hu, Renshuai Tao, Xianglong Liu

Abstract

There is a growing concern about adversarial attacks against automatic speech recognition (ASR) systems. Although research into targeted universal adversarial examples (AEs) has progressed, current methods are constrained by inefficient exploitation of audio features, demonstrating insufficient attack ability and robustness in the physical world. To solve this problem, we propose a phoneme-tailored attack (PTA) to improve the quality of the generated AEs. Specifically, to improve attack ability, we propose a Diverse Audio Composition Enrichment method, which enhances the utilization of audio features through phoneme-level slicing and recombination. To adapt AEs to complex environments, we propose a Natural Noise Pattern Guidance method to align AEs with natural noise patterns to improve their robustness. Experiments show that our method achieves an average accuracy of more than 72.34% and 98% with and without a norm constraint, and also demonstrates excellent performance in terms of generalization across datasets and resilience to MP3 compression.

BibTeX
@inproceedings{icassp2025_generatingtarget,
  title = {Generating Targeted Universal Adversarial Perturbation against Automatic Speech Recognition via Phoneme Tailoring},
  author = {Yujun Zhang and Yanqu Chen and Jiakai Wang and Jin Hu and Renshuai Tao and Xianglong Liu},
  booktitle = {ICASSP 2025},
  year = {2025}
}