ICASSP 2016accepted0 citations

Geo-location dependent deep neural network acoustic model for speech recognition

Guoli Ye, Chaojun Liu, Yifan Gong

Abstract

Users from the same geo-location region exhibit similar acoustic characteristics, e.g., they have similar accent; even more, they may have similar preference to device. In this paper, we propose to build geo-location dependent deep neural network for speech recognition, where the geo-location signal is inferred from users' GPS. During runtime, the server will base on a user's geo-location to select the right model to recognize his voice. We tackle three major issues associated with this model: high train/deployment cost, large model size, and train data sparsity. Our solution is featured by its low cost, thus practical for production modeling. We also discuss the reliability of GPS signal in practical use. The proposed model is evaluated on Microsoft Chinese voice search and Cortana live test set. Among 12 provinces, it shows an overall 4.8% relative character error rate reduction, over a strong baseline production-level model, with only 50% model size increase. The gain is larger for the low-resource provinces, with relative error rate reduction up to 9%.

BibTeX
@inproceedings{icassp2016_geolocationdepen,
  title = {Geo-location dependent deep neural network acoustic model for speech recognition},
  author = {Guoli Ye and Chaojun Liu and Yifan Gong},
  booktitle = {ICASSP 2016},
  year = {2016}
}
Geo-location dependent deep neural network acoustic model for speech recognition · ICASSP 2016