Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
As the harm caused by fake news grows, the task of detecting and grounding multi-modal media manipulation (DGM4) is gaining more attention. Existing multimodal methods overlook fine-grained semantic alignment between visual and textual modalities, thereby limiting their ability to detect sophisticat