2019
Multimodal Speaker Adaptation of Acoustic Model and Language Model for Asr Using Speaker Face Embedding
ICASSP 2019accepted
We present an investigation into the adaptation of the acoustic model and the language model for automatic speech recognition (ASR) using speaker face for transcription of a multimedia dataset. We begin by overviewing relevant previous work on the integration of visual signals into ASR systems. Our…