Analysis of DNN approaches to speaker identification
Pavel Matejka, Ondrej Glembek, Ondrej Novotný, Oldrich Plchot, Frantisek Grézl, Lukás Burget, Jan Cernocký
Abstract
This work studies the usage of the Deep Neural Network (DNN) Bottleneck (BN) features together with the traditional MFCC features in the task of i-vector-based speaker recognition. We decouple the sufficient statistics extraction by using separate GMM models for frame alignment, and for statistics normalization and we analyze the usage of BN and MFCC features (and their concatenation) in the two stages. We also show the effect of using full-covariance GMM models, and, as a contrast, we compare the result to the recent DNN-alignment approach. On the NIST SRE2010, telephone condition, we show 60% relative gain over the traditional MFCC baseline for EER (and similar for the NIST DCF metrics), resulting in 0.94% EER.
BibTeX
@inproceedings{icassp2016_analysisofdnnapp,
title = {Analysis of DNN approaches to speaker identification},
author = {Pavel Matejka and Ondrej Glembek and Ondrej Novotný and Oldrich Plchot and Frantisek Grézl and Lukás Burget and Jan Cernocký},
booktitle = {ICASSP 2016},
year = {2016}
}