SpecViT: A Custom Vision-Transformer based Approach for Audio Deepfake Detection
Sharmistha Modak, Arnab Kumar Das, Ruchira Naskar
Abstract
Degrees of hyper-realism already attained by present-day deepfake technology poses one of the biggest social threats of today. Deepfakes may involve multiple forms of media including audio, video and images. While most of the literature deals with threats posed by visual synthetic media, the endeavor to uncover audio deepfakes is still evolving; it demands more extensive investigation. Our attempt towards audio deepfake detection in this article, involves investigating spectral patterns present in audio spectrograms, captured appropriately by a two-attention vision transformer model (SpecViT2A), finally exploited for discriminating synthetic audios from pristine ones. We employed our model for successful identification of audio as well as multimodal deepfakes, yielding best accuracy over 99%, F1-score 0.9911, and EER as low as 3.5 on ASVSpoof 2021.
BibTeX
@inproceedings{icassp2025_specvitacustomvi,
title = {SpecViT: A Custom Vision-Transformer based Approach for Audio Deepfake Detection},
author = {Sharmistha Modak and Arnab Kumar Das and Ruchira Naskar},
booktitle = {ICASSP 2025},
year = {2025}
}