ICASSP 2022accepted0 citations

Improving End-To-End Speech Translation Model with Bert-Based Contextual Information

Jeong-Uk Bang, Min-Kyu Lee, Seung Yun, Sang-Hun Kim

Abstract

This paper proposes an end-to-end speech translation system that utilizes contextual information. Contextual information helps clarify the meaning of the utterances. However, conventional end-to-end speech translation (E2E-ST) is primarily designed to handle single-utterance. Thus, we introduce a context encoder that extracts contextual information from previous translation results. Here, the context encoder obtains high-quality contextual information by adopting the BERT model. Then, we combine it with speech information extracted from speech signals to generate translation results. On the widely used TED-based speech translation corpus, we show that the results of the contextual E2E-ST model are significantly better than those of the single utterance-based E2E-ST model. Furthermore, we demonstrate that contextual information contributes to the processing of unclearly spoken utterances as well as ambiguity caused by pronouns and homophones.

BibTeX
@inproceedings{icassp2022_improvingendtoen,
  title = {Improving End-To-End Speech Translation Model with Bert-Based Contextual Information},
  author = {Jeong-Uk Bang and Min-Kyu Lee and Seung Yun and Sang-Hun Kim},
  booktitle = {ICASSP 2022},
  year = {2022}
}