ICASSP 2025accepted0 citations

PET: High-Frequency Temporal Self-Consistency Learning for Partially Deepfake Audio Localization

Jiayi He, Jiangyan Yi, Jianhua Tao, Siding Zeng

Abstract

Partially deepfake audio attacks have attracted the attention recently, and the demand for locating the manipulation regions of partially deepfake audio arises accordingly. However, existing methods are usually proposed based on frame-level authenticity detection or splicing boundaries detection, neglecting the temporal self-consistency of audio. In this paper, we propose a novel method for partially deepfake audio localization based on temporal self-consistency learning via high-frequency components, named as PET. The results demonstrates that, in ADD 2023 Track 2 eval set, it could achieve the segment F1-score at 0.7397 without any data augmentation strategies, which is 21.94% higher than that of the system ranked 1st on the leaderboard. It also confirms the effectiveness and well generalization ability of PET.

BibTeX
@inproceedings{icassp2025_pethighfrequency,
  title = {PET: High-Frequency Temporal Self-Consistency Learning for Partially Deepfake Audio Localization},
  author = {Jiayi He and Jiangyan Yi and Jianhua Tao and Siding Zeng},
  booktitle = {ICASSP 2025},
  year = {2025}
}
PET: High-Frequency Temporal Self-Consistency Learning for Partially Deepfake Audio Localization · ICASSP 2025