Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding
Abstract
In response to the security risks posed by the realistic propagation of manipulated media data, detecting and grounding multi-modal media manipulation has received attention as a challenging task. However, there is multi-modal contribution imbalance on current approach for cross-modal learning, which affects model performance optimisation. To this end, we propose an Adaptive Contribution Modulation (ACM) framework to solve the problem of multi-modal contribution imbalance. To balance the image and text embedding features before fusion, we propose adaptive weight decision to computes dynamic weights for fusion features, which enable more adaptive and robust decision-making. Meanwhile, we propose contribution modulation block, which dynamically governs the contributions of different modalities for optimization. Based on cross-modal contrastive learning, we balance image and text embeddings contribution through multi-modal contribution balanced learning, which makes better use of the semantic correlation of all modalities. We conduct experiments on the DGM4 dataset, which demonstrate the superior performance of our approach through compared to state-of-the-art methods.
BibTeX
@inproceedings{icassp2025_adaptivecontribu,
title = {Adaptive Contribution Modulation For Multi-Modal Manipulation Media Detection and Grounding},
author = {Yixiang Li and Biao Leng},
booktitle = {ICASSP 2025},
year = {2025}
}