Speech-to-Speech Translation for a Real-world Unwritten Language
Peng-Jen Chen, Kevin Tran, Yilin Yang, Jingfei Du, Justine Kao, Yu-An Chung, Paden Tomasello, Paul-Ambroise Duquenne
Abstract
We study speech-to-speech translation (S2ST) that translates speech from one language into another language and focuses on building systems to support languages without standard text writing systems. We use English-Taiwanese Hokkien as a case study, and present an end-to-end solution from training data collection, modeling choices to benchmark dataset release. First, we present efforts on creating human annotated data, automatically mining data from large unlabeled speech datasets, and adopting pseudo-labeling to produce weakly supervised data. On the modeling, we take advantage of recent advances in applying self-supervised discrete representations as target for prediction in S2ST and show the effectiveness of leveraging additional text supervision from Mandarin, a language similar to Hokkien, in model training. Finally, we release an S2ST benchmark set to facilitate future research in this field.
BibTeX
@inproceedings{chen-etal-2023-speech,
title = "Speech-to-Speech Translation for a Real-world Unwritten Language",
author = "Chen, Peng-Jen and
Tran, Kevin and
Yang, Yilin and
Du, Jingfei and
Kao, Justine and
Chung, Yu-An and
Tomasello, Paden and
Duquenne, Paul-Ambroise and
Schwenk, Holger and
Gong, Hongyu and
Inaguma, Hirofumi and
Popuri, Sravya and
Wang, Changhan and
Pino, Juan and
Hsu, Wei-Ning and
Lee, Ann",
editor = "Rogers, Anna and
Boyd-Graber, Jordan and
Okazaki, Naoaki",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2023",
month = jul,
year = "2023",
address = "Toronto, Canada",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2023.findings-acl.307/",
doi = "10.18653/v1/2023.findings-acl.307",
pages = "4969--4983"
}