ICASSP 2020accepted0 citations

Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker Embeddings

Aku Rouhe, Tuomas Kaseva, Mikko Kurimo

Abstract

In speaker-aware training, a speaker embedding is appended to DNN input features. This allows the DNN to effectively learn representations, which are robust to speaker variability.We apply speaker-aware training to attention-based end-to-end speech recognition. We show that it can improve over a purely end-to-end baseline. We also propose speaker-aware training as a viable method to leverage untranscribed, speaker annotated data.We apply state-of-the-art embedding approaches, both i-vectors and neural embeddings, such as x-vectors. We experiment with embeddings trained in two conditions: on the fixed ASR data, and on a large untranscribed dataset. We run our experiments on the TED-LIUM and Wall Street Journal datasets. No embedding consistently outperforms all others, but in many settings neural embeddings outperform i-vectors.

BibTeX
@inproceedings{icassp2020_speakerawaretrai,
  title = {Speaker-Aware Training of Attention-Based End-to-End Speech Recognition Using Neural Speaker Embeddings},
  author = {Aku Rouhe and Tuomas Kaseva and Mikko Kurimo},
  booktitle = {ICASSP 2020},
  year = {2020}
}