Improving Identification of System-Directed Speech Utterances by Deep Learning of ASR-Based Word Embeddings and Confidence Metrics
Vilayphone Vilaysouk, Amr Nour-Eldin, Dermot Connolly
Abstract
In this paper, we extend our previous work on the detection of system-directed speech utterances. This type of binary classification can be used by virtual assistants to create a more natural and fluid interaction between the system and the user. We explore two methods that both improve the Equal-Error-Rate (EER) performance of the previous model. The first exploits the supplementary information independently captured by ASR models through integrating ASR decoder-based features as additional inputs to the final classification stage of the model. This relatively improves EER performance by 13%. The second proposed method further integrates word embeddings into the architecture and, when combined with the first method, achieves a significant EER performance improvement of 48%, relative to that of the baseline.
BibTeX
@inproceedings{icassp2021_improvingidentif,
title = {Improving Identification of System-Directed Speech Utterances by Deep Learning of ASR-Based Word Embeddings and Confidence Metrics},
author = {Vilayphone Vilaysouk and Amr Nour-Eldin and Dermot Connolly},
booktitle = {ICASSP 2021},
year = {2021}
}