ICASSP 2022accepted0 citations

Temporal Early Exiting for Streaming Speech Commands Recognition

Raphael Tang, Karun Kumar, Ji Xin, Piyush Vyas, Wenyan Li, Gefei Yang, Yajie Mao, G. Craig Murray

Abstract

Limited-vocabulary speech commands recognition is the task of classifying a short utterance as one of several speech commands, for which neural networks obtain state-of-the-art results. In particular, recurrent neural networks represent a common approach for streaming commands recognition systems. In this paper, we explore resource-efficient methods to short-circuit such systems in the time domain when the model is confident in its prediction. We propose applying a frame-level labeling objective to further improve the efficiency–accuracy trade-off. On two datasets in limited-vocabulary commands recognition, our best method achieves an average time savings of 45% of the utterance without reducing the absolute accuracy by more than 0.6 points. We show that the per-instance savings depend on the length of the unique prefix in the phonemes across a dataset.

BibTeX
@inproceedings{icassp2022_temporalearlyexi,
  title = {Temporal Early Exiting for Streaming Speech Commands Recognition},
  author = {Raphael Tang and Karun Kumar and Ji Xin and Piyush Vyas and Wenyan Li and Gefei Yang and Yajie Mao and G. Craig Murray and Jimmy Lin},
  booktitle = {ICASSP 2022},
  year = {2022}
}