ICASSP 2025accepted0 citations

From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background

Zhaolin Wei, Dengpan Ye, Jiacheng Deng, Yuhan Lin

Abstract

Music deepfake detection is aimed at identifying whether songs are generated by AI. Current methods usually separate vocals from background music for detection, but this could leave residual forgery information in the background. Our study demonstrates for the first time that incorporating background forgery information with vocals can improve detection accuracy. Furthermore, our findings show that using background features can reduce EER by an average of about 2% on existing frameworks. Based on this observation, we propose a novel Hybrid Frontend that captures generalized features from both vocal and background music. The Hybrid Frontend comprises two branches: vocal and background music part. Specifically, the vocal part uses the Sinconv encoder as a deeply embedded feature extractor. The latter captures background variation by fine-tuning the pre-trained model with adapters. Experimental results demonstrate that our method outperforms the vocal-only detection on WildSVDD dataset, achieving an EER of 8.53%, which is 1.3% lower.

BibTeX
@inproceedings{icassp2025_fromvoicestobeat,
  title = {From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background},
  author = {Zhaolin Wei and Dengpan Ye and Jiacheng Deng and Yuhan Lin},
  booktitle = {ICASSP 2025},
  year = {2025}
}
From Voices to Beats: Enhancing Music Deepfake Detection by Identifying Forgeries in Background · ICASSP 2025