ACL 2024long5 citations

Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup

Xuxin Cheng, Ziyu Yao, Yifei Xin, Hao An, Hongxiang Li, Yaowei Li, Yuexian Zou

Abstract

Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently. It has been verified that visual information brings greater performance gains when the textual information is limited. However, most previous works ignore to take advantage of the complete textual inputs and the limited textual inputs at the same time, which limits the overall performance. To solve this issue, we propose a mixup method termed Soul-Mix to enhance MMT by using visual information more effectively. We mix the predicted translations of complete textual input and the limited textual inputs. Experimental results on the Multi30K dataset of three translation directions show that our Soul-Mix significantly outperforms existing approaches and achieves new state-of-the-art performance with fewer parameters than some previous models. Besides, the strength of Soul-Mix is more obvious on more challenging MSCOCO dataset which includes more out-of-domain instances with lots of ambiguous verbs.

BibTeX
@inproceedings{cheng-etal-2024-soul,
    title = "Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup",
    author = "Cheng, Xuxin  and
      Yao, Ziyu  and
      Xin, Yifei  and
      An, Hao  and
      Li, Hongxiang  and
      Li, Yaowei  and
      Zou, Yuexian",
    editor = "Ku, Lun-Wei  and
      Martins, Andre  and
      Srikumar, Vivek",
    booktitle = "Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = aug,
    year = "2024",
    address = "Bangkok, Thailand",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.acl-long.608/",
    doi = "10.18653/v1/2024.acl-long.608",
    pages = "11283--11294"
}
Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup · ACL 2024