EMNLP 2024main1 citations

Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair

Yusuke Sakai, Mana Makinae, Hidetaka Kamigaito, Taro Watanabe

Abstract

In Simultaneous Machine Translation (SiMT), training with a simultaneous interpretation (SI) corpus is an effective method for achieving high-quality yet low-latency. However, constructing such a corpus is challenging due to high costs, and limitations in annotator capabilities, and as a result, existing SI corpora are limited. Therefore, we propose a method to convert existing speech translation (ST) corpora into interpretation-style corpora, maintaining the original word order and preserving the entire source content using Large Language Models (LLM-SI-Corpus). We demonstrate that fine-tuning SiMT models using the LLM-SI-Corpus reduces latency while achieving better quality compared to models fine-tuned with other corpora in both speech-to-text and text-to-text settings. The LLM-SI-Corpus is available at https://github.com/yusuke1997/LLM-SI-Corpus.

BibTeX
@inproceedings{sakai-etal-2024-simultaneous,
    title = "Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair",
    author = "Sakai, Yusuke  and
      Makinae, Mana  and
      Kamigaito, Hidetaka  and
      Watanabe, Taro",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.1248/",
    doi = "10.18653/v1/2024.emnlp-main.1248",
    pages = "22375--22398"
}
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair · EMNLP 2024