ICASSP 2016accepted0 citations

Deep neural network based posteriors for text-dependent speaker verification

Subhadeep Dey, Srikanth R. Madikeri, Marc Ferras, Petr Motlícek

Abstract

The i-vector and Joint Factor Analysis (JFA) systems for text-dependent speaker verification use sufficient statistics computed from a speech utterance to estimate speaker models. These statistics average the acoustic information over the utterance thereby losing all the sequence information. In this paper, we study explicit content matching using Dynamic Time Warping (DTW) and present the best achievable error rates for speaker-dependent and speaker-independent content matching. For this purpose, a Deep Neural Network/Hidden Markov Model Automatic Speech Recognition (DNN/HMM ASR) system is used to extract content-related posterior probabilities. This approach outperforms systems using Gaussian mixture model posteriors by at least 50% Equal Error Rate (EER) on the RSR2015 in content mismatch trials. DNN posteriors are also used in i-vector and JFA systems, obtaining EERs as low as 0.02%.

BibTeX
@inproceedings{icassp2016_deepneuralnetwor,
  title = {Deep neural network based posteriors for text-dependent speaker verification},
  author = {Subhadeep Dey and Srikanth R. Madikeri and Marc Ferras and Petr Motlícek},
  booktitle = {ICASSP 2016},
  year = {2016}
}