ICASSP 2024accepted0 citations

End-to-End Speech Translation with Mutual Knowledge Distillation

Hao Wang, Zhengshan Xue, Yikun Lei, Deyi Xiong

Abstract

Multi-task learning (MTL) is widely used to improve end-to-end speech translation (ST), which implicitly transfer knowledge from auxiliary automatic speech recognition (ASR) and/or machine translation (MT) to ST through shared modules. In this study, we find that triple-task MTL (ST+MT+ASR) suffers from a knowledge transfer limitation that leads to performance stagnation compared with dual-task MTL (ST+MT or ST+ASR). To address this issue, we propose a simple yet effective method, ST-MKD (Speech Translation with Mutual Knowledge Distillation). In ST-MKD, we employ a mutual knowledge distillation framework to mutually enhance dual-task MTL models with different knowledge bases, and explore regularization to maintain the consistency of the task representations. Experiments on the ST benchmark dataset MuST-C show that ST-MKD significantly outperforms strong MTL baseline and achieves state-of-the-art performance under three speech pre-training settings. Further analyses confirm that our approach effectively overcomes the knowledge transfer limitation of triple-task MTL.

BibTeX
@inproceedings{icassp2024_endtoendspeechtr,
  title = {End-to-End Speech Translation with Mutual Knowledge Distillation},
  author = {Hao Wang and Zhengshan Xue and Yikun Lei and Deyi Xiong},
  booktitle = {ICASSP 2024},
  year = {2024}
}
End-to-End Speech Translation with Mutual Knowledge Distillation · ICASSP 2024