ACL 2025long0 citations

Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation

Andong Chen, Yuchen Song, Kehai Chen, Xuefeng Bai, Muyun Yang, Liqiang Nie, Jie Liu, Tiejun Zhao

Abstract

Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we introduce a stable diffusion-based imagination network into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing the multimodel MT. Particularly, we build heuristic feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of visual information, which breaks the high-cost bottleneck of image annotation in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 14 BLEU points on Multi30K and MSCOCO multimodal MT benchmarks.

BibTeX
@inproceedings{chen-etal-2025-make,
    title = "Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation",
    author = "Chen, Andong  and
      Song, Yuchen  and
      Chen, Kehai  and
      Bai, Xuefeng  and
      Yang, Muyun  and
      Nie, Liqiang  and
      Liu, Jie  and
      Zhao, Tiejun  and
      Zhang, Min",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.acl-long.1289/",
    doi = "10.18653/v1/2025.acl-long.1289",
    pages = "26567--26583",
    ISBN = "979-8-89176-251-0"
}
Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation · ACL 2025