ACL 2023long34 citations

SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations

Paul-Ambroise Duquenne, Hongyu Gong, Ning Dong, Jingfei Du, Ann Lee, Vedanuj Goswami, Changhan Wang, Juan Pino

Abstract

We present SpeechMatrix, a large-scale multilingual corpus of speech-to-speech translations mined from real speech of European Parliament recordings. It contains speech alignments in 136 language pairs with a total of 418 thousand hours of speech. To evaluate the quality of this parallel speech, we train bilingual speech-to-speech translation models on mined data only and establish extensive baseline results on EuroParl-ST, VoxPopuli and FLEURS test sets. Enabled by the multilinguality of SpeechMatrix, we also explore multilingual speech-to-speech translation, a topic which was addressed by few other works. We also demonstrate that model pre-training and sparse scaling using Mixture-of-Experts bring large gains to translation performance. The mined data and models will be publicly released

BibTeX
@inproceedings{duquenne-etal-2023-speechmatrix,
    title = "{S}peech{M}atrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations",
    author = "Duquenne, Paul-Ambroise  and
      Gong, Hongyu  and
      Dong, Ning  and
      Du, Jingfei  and
      Lee, Ann  and
      Goswami, Vedanuj  and
      Wang, Changhan  and
      Pino, Juan  and
      Sagot, Beno{\^i}t  and
      Schwenk, Holger",
    editor = "Rogers, Anna  and
      Boyd-Graber, Jordan  and
      Okazaki, Naoaki",
    booktitle = "Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)",
    month = jul,
    year = "2023",
    address = "Toronto, Canada",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2023.acl-long.899/",
    doi = "10.18653/v1/2023.acl-long.899",
    pages = "16251--16269"
}
SpeechMatrix: A Large-Scale Mined Corpus of Multilingual Speech-to-Speech Translations · ACL 2023