Fooling the Forgers: A Multi-Stage Framework for Audio Deepfake Detection
Gautam Siddharth Kashyap, Zohaib Hasan Siddiqui, Mohammad Anas Azeez, Rafiq Ali, Shantanu Kumar, Navin Kamuni, Jiechao Gao
Abstract
Audio deepfakes represent a risk to society as they can deteriorate society’s trust in any audio. In this paper, we present a novel approach for audio deepfake detection using Generative Adversarial Networks (GANs) and contrastive learning in a multi-stage detection framework. In our process, we apply the Pre-trained Models (PTM) to extract all suitable audio phonetics, speaker identity, and other spatial prosodic features or contents, which are crucial for the model. We enhance the model’s performance by utilizing a GAN data augmentation strategy in combination with HiFi-GAN. The Contrastive learning approach is then used for improving the model’s ability to discriminate real speech from fake speech. Our experiments demonstrate that this method is superior to existing methodologies in detection and robustness.
BibTeX
@inproceedings{icassp2025_foolingtheforger,
title = {Fooling the Forgers: A Multi-Stage Framework for Audio Deepfake Detection},
author = {Gautam Siddharth Kashyap and Zohaib Hasan Siddiqui and Mohammad Anas Azeez and Rafiq Ali and Shantanu Kumar and Navin Kamuni and Jiechao Gao},
booktitle = {ICASSP 2025},
year = {2025}
}