ICASSP 2023accepted0 citations

Understanding Shared Speech-Text Representations

Gary Wang, Kyle Kastner, Ankur Bapna, Zhehuai Chen, Andrew Rosenberg, Bhuvana Ramabhadran, Yu Zhang

Abstract

Recently, a number of approaches to train speech models by incorporating text into end-to-end models have been developed, with Maestro advancing state-of-the-art automatic speech recognition (ASR) and Speech Translation (ST) performance. In this paper, we expand our understanding of the resulting shared speech-text representations with two types of analyses. First we examine the limits of speech-free domain adaptation, finding that a corpus-specific duration model for speech-text alignment is the most important component for learning a shared speech-text representation. Second, we inspect the similarities between activations of unimodal (speech or text) encoders as compared to the activations of a shared encoder. We find that the shared encoder learns a more compact and overlapping speech-text representation than the uni-modal encoders. We hypothesize that this partially explains the effectiveness of the Maestro shared speech-text representations.

BibTeX
@inproceedings{icassp2023_understandingsha,
  title = {Understanding Shared Speech-Text Representations},
  author = {Gary Wang and Kyle Kastner and Ankur Bapna and Zhehuai Chen and Andrew Rosenberg and Bhuvana Ramabhadran and Yu Zhang},
  booktitle = {ICASSP 2023},
  year = {2023}
}