COLING 2025main0 citations

Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation

Bo Xu, Haiqi Jiang, Jie Wei, Hongyu Jing, Ming Du, Hui Song, Hongya Wang, Yanghua Xiao

Abstract

Multimodal named entity recognition (MNER) extends traditional named entity recognition (NER) by integrating visual and textual information. However, current methods still face significant challenges due to the text-image mismatch problem. Recent advancements in text-to-image synthesis provide promising solutions, as synthesized images can introduce additional visual context to enhance MNER model performance. To fully leverage the benefits of both original and synthesized images, we propose an adaptive mixup image augmentation method. This method generates augmented images by determining the mixing ratio based on the matching score between the text and image, utilizing a triplet loss-based Gaussian Mixture Model (TL-GMM). Our approach is highly adaptable and can be seamlessly integrated into existing MNER models. Extensive experiments demonstrate consistent performance improvements, and detailed ablation studies and case studies confirm the effectiveness of our method.

BibTeX
@inproceedings{xu-etal-2025-enhancing,
    title = "Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation",
    author = "Xu, Bo  and
      Jiang, Haiqi  and
      Wei, Jie  and
      Jing, Hongyu  and
      Du, Ming  and
      Song, Hui  and
      Wang, Hongya  and
      Xiao, Yanghua",
    editor = "Rambow, Owen  and
      Wanner, Leo  and
      Apidianaki, Marianna  and
      Al-Khalifa, Hend  and
      Eugenio, Barbara Di  and
      Schockaert, Steven",
    booktitle = "Proceedings of the 31st International Conference on Computational Linguistics",
    month = jan,
    year = "2025",
    address = "Abu Dhabi, UAE",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.coling-main.122/",
    pages = "1802--1812"
}
Enhancing Multimodal Named Entity Recognition through Adaptive Mixup Image Augmentation · COLING 2025