ICASSP 2024accepted0 citations

Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation

Weichen Zhao, Yuxing Lu, Ge Jiao, Yuan Yang

Abstract

Detecting and Grounding Multi-Modal Media Manipulation (DGM <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">4</sup> ) is an emerging task that aims to identify and locate manipulated elements in both textual and visual media. Given the complexity of this task, the model requires more sophisticated reasoning capabilities to align multi-modal features and capture forgery traces. To this end, we propose a Concentrated reasoning and Unified reconstruction framework (CrUr) for DGM <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">4</sup> . Instead of adhering to traditional hierarchical reasoning paradigms, we directly carry out all inference tasks using integrated multi-modal features. Specifically, we extract and align features at a finer granularity, capturing subtle differences that may indicate manipulation by leveraging advanced mask signal modeling. Moreover, to adapt to fine-grained reasoning tasks, we design a transformer-based Reconstruction Harmonizer to facilitate more complex interactions among the reconstructed features, ultimately obtaining integrated features. Experimental results on the DGM <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">4</sup> datasets show that our method achieves state-of-the-art performances.

BibTeX
@inproceedings{icassp2024_concentratedreas,
  title = {Concentrated Reasoning and Unified Reconstruction for Multi-Modal Media Manipulation},
  author = {Weichen Zhao and Yuxing Lu and Ge Jiao and Yuan Yang},
  booktitle = {ICASSP 2024},
  year = {2024}
}