EMNLP 2022main21 citations

T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation

Paul-Ambroise Duquenne, Hongyu Gong, Benoît Sagot, Holger Schwenk

Abstract

We present a new approach to perform zero-shot cross-modal transfer between speech and text for translation tasks. Multilingual speech and text are encoded in a joint fixed-size representation space. Then, we compare different approaches to decode these multimodal and multilingual fixed-size representations, enabling zero-shot translation between languages and modalities. All our models are trained without the need of cross-modal labeled translation data.Despite a fixed-size representation, we achieve very competitive results on several text and speech translation tasks. In particular, we significantly improve the state-of-the-art for zero-shot speech translation on Must-C. Incorporating a speech decoder in our framework, we introduce the first results for zero-shot direct speech-to-speech and text-to-speech translation.

BibTeX
@inproceedings{duquenne-etal-2022-modules,
    title = "{T}-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation",
    author = "Duquenne, Paul-Ambroise  and
      Gong, Hongyu  and
      Sagot, Beno{\^i}t  and
      Schwenk, Holger",
    editor = "Goldberg, Yoav  and
      Kozareva, Zornitsa  and
      Zhang, Yue",
    booktitle = "Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing",
    month = dec,
    year = "2022",
    address = "Abu Dhabi, United Arab Emirates",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.emnlp-main.391/",
    doi = "10.18653/v1/2022.emnlp-main.391",
    pages = "5794--5806"
}
T-Modules: Translation Modules for Zero-Shot Cross-Modal Machine Translation · EMNLP 2022