ICASSP 2025accepted0 citations

What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting

Yixuan Xiao, Ngoc Thang Vu

Abstract

The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers’ training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers’ training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems.

BibTeX
@inproceedings{icassp2025_whataffectsthepe,
  title = {What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting},
  author = {Yixuan Xiao and Ngoc Thang Vu},
  booktitle = {ICASSP 2025},
  year = {2025}
}
What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting · ICASSP 2025