ICASSP 2019accepted0 citations

Multimodal Speaker Adaptation of Acoustic Model and Language Model for Asr Using Speaker Face Embedding

Yasufumi Moriya, Gareth J. F. Jones

Abstract

We present an investigation into the adaptation of the acoustic model and the language model for automatic speech recognition (ASR) using speaker face for transcription of a multimedia dataset. We begin by overviewing relevant previous work on the integration of visual signals into ASR systems. Our experimental investigation shows a small improvement in word error rate (WER) for the transcription of a collection of instruction videos using adaptation of the acoustic model and the language model with fixed-length face embedding vectors. We also present potential approaches to integrating human facial information, and body gestures into ASR as further directions for research on this topic.

BibTeX
@inproceedings{icassp2019_multimodalspeake,
  title = {Multimodal Speaker Adaptation of Acoustic Model and Language Model for Asr Using Speaker Face Embedding},
  author = {Yasufumi Moriya and Gareth J. F. Jones},
  booktitle = {ICASSP 2019},
  year = {2019}
}