TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLU
Anastasios Alexandridis, Kanthashree Mysore Sathyendra, Grant P. Strimel, Pavel Kveton, Jon Webb, Athanasios Mouchtaris
Abstract
On-device spoken language understanding (SLU) offers the potential for significant latency savings compared to cloud-based processing, as the audio stream does not need to be transmitted to a server. We present Tiny Signal-to-interpretation (TinyS2I), an end-to-end on-device SLU approach which is focused on heavily resource constrained devices. TinyS2I brings latency reduction without accuracy degradation, by exploiting use cases when the distribution of utterances that users speak to a device is largely heavy-tailed. The model is tailored to process on-device frequent utterances with support for dynamic contextual content, while deferring all other requests to the cloud. Compared to a powerful baseline, we demonstrate that TinyS2I achieves comparable performance, while offering latency gains due to local processing.
BibTeX
@inproceedings{icassp2022_tinys2iasmallfoo,
title = {TINYS2I: A Small-Footprint Utterance Classification Model with Contextual Support for On-Device SLU},
author = {Anastasios Alexandridis and Kanthashree Mysore Sathyendra and Grant P. Strimel and Pavel Kveton and Jon Webb and Athanasios Mouchtaris},
booktitle = {ICASSP 2022},
year = {2022}
}