ICASSP 2025accepted0 citations

Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection

Yunqi Hao, Minqiang Xu, Yihao Chen, Yanyan Liu, Liang He, Lei Fang, Lin Liu

Abstract

Audio deepfake refers to the technology of synthesizing speech using deep learning or large model algorithms. Compared to human voice, synthetic deepfake speech exhibits artifacts at global and local levels, which can be leveraged by audio deepfake detection (ADD) to distinguish real and fake speech. In this paper, we designed the spectro-temporal cross aggregation (STCA) module and the local multi-scale dynamic convolution (LMDC) module to extract global and local artifacts for detecting forged information, respectively. The STCA module utilizes a dual-branch structure with cross-attention, extracting global temporal and frequency features through the branches and aggregating mutual artifacts via cross-attention. The LMDC module uses multi-scale dynamic convolution for grouped features, to extract local information. Experimental results on multiple test sets demonstrate the effectiveness of our method. Specifically, we achieved an EER of 1.87% on the ASVspoof2021 DF evaluation set, surpassing the current state-of-the-art system by a relative 14.6%.

BibTeX
@inproceedings{icassp2025_integratingspect,
  title = {Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection},
  author = {Yunqi Hao and Minqiang Xu and Yihao Chen and Yanyan Liu and Liang He and Lei Fang and Lin Liu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
Integrating Spectro-Temporal Cross Aggregation and Multi-Scale Dynamic Learning for Audio Deepfake Detection · ICASSP 2025