Recent Advances in Direct Speech-to-text Translation
Chen Xu, Rong Ye, Qianqian Dong, Chengqi Zhao, Tom Ko, Mingxuan Wang, Tong Xiao, Jingbo Zhu
Abstract
Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech translation aiming to summarize the current state-of-the-art techniques. First, we categorize the existing research work into three directions based on the main challenges --- modeling burden, data scarcity, and application issues. To tackle the problem of modeling burden, two main structures have been proposed, encoder-decoder framework (Transformer and the variants) and multitask frameworks. For the challenge of data scarcity, recent work resorts to many sophisticated techniques, such as data augmentation, pre-training, knowledge distillation, and multilingual modeling. We analyze and summarize the application issues, which include real-time, segmentation, named entity, gender bias, and code-switching. Finally, we discuss some promising directions for future work.
BibTeX
@inproceedings{ijcai2023p761,
title = {Recent Advances in Direct Speech-to-text Translation},
author = {Xu, Chen and Ye, Rong and Dong, Qianqian and Zhao, Chengqi and Ko, Tom and Wang, Mingxuan and Xiao, Tong and Zhu, Jingbo},
booktitle = {Proceedings of the Thirty-Second International Joint Conference on
Artificial Intelligence, {IJCAI-23}},
publisher = {International Joint Conferences on Artificial Intelligence Organization},
editor = {Edith Elkind},
pages = {6796--6804},
year = {2023},
month = {8},
note = {Survey Track},
doi = {10.24963/ijcai.2023/761},
url = {https://doi.org/10.24963/ijcai.2023/761},
}