ICASSP 2025accepted0 citations

Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection

Siyu Liu, Zhida Zhang, Junxian Duan, Jie Cao, Aihua Zheng

Abstract

With the advancement of generative AI, distinguishing real and AI-generated faces in videos has become increasingly challenging. However, traditional methods struggle to capture local details and temporal dynamics simultaneously, making it difficult to achieve high detection accuracy while maintaining low computational overhead. To address this problem, we propose a Dual-branch SpatioTemporal-Planar Network (Dual-PST) based on the selective state-space model. It is capable of extracting image features and temporal relations simultaneously, while maintaining linear computational consumption. Specifically, we design a Multi-Selective State-Space module (MS3) that can extract global features from image typography consisting of consecutive video frames by scanning them in multiple sequences. To further enhance temporal modeling capabilities, we propose a Sequential Tri-frame Local module, which captures inter-frame temporal relationships and local features by temporally splicing single-frame features. These features are first extracted using MS3 and then further enhanced through inter-frame masking operations. Experimental results show that Dual-PST significantly improves detection accuracy while maintaining low computational complexity and strong model robustness.

BibTeX
@inproceedings{icassp2025_dualpstdualbranc,
  title = {Dual-PST: Dual-Branch SpatioTemporal-Planar Network for Video Forgery Detection},
  author = {Siyu Liu and Zhida Zhang and Junxian Duan and Jie Cao and Aihua Zheng},
  booktitle = {ICASSP 2025},
  year = {2025}
}