ICASSP 2019accepted0 citations

Formant-gaps Features for Speaker Verification Using Whispered Speech

Abinay Reddy Naini, M. V. Achuth Rao, Prasanta Kumar Ghosh

Abstract

In this work, we propose a new feature based on formants for whispered speaker verification (SV) task, where neutral data is used for enrollment and whispered recordings are used for test. Such a mismatch between enrollment and test often degrades the performance of whispered SV systems due to the difference in acoustic characteristics of whispered and neutral speech. We hypothesize that the proposed formant and formant gap (F oG) features are more invariant to the modes of speech in capturing speaker specific information compared to traditional baseline features for SV including mel frequency cepstral coefficients (MFCC) and auditory-inspired amplitude modulation features (AAMF). Whispered SV experiments with 714 speakers comprising 29232 neutral and 22932 whispered recordings reveal that the equal error rate (EER) using the proposed features is lower than that using the best baseline features by ~3.79% (absolute). It was also observed that at least four whispered recordings during enrollment are required for the baseline features to perform at par with the proposed features. However, it was found that the best performing baseline features yield an EER for neutral SV task which is ~1.88% higher than that using the proposed features.

BibTeX
@inproceedings{icassp2019_formantgapsfeatu,
  title = {Formant-gaps Features for Speaker Verification Using Whispered Speech},
  author = {Abinay Reddy Naini and M. V. Achuth Rao and Prasanta Kumar Ghosh},
  booktitle = {ICASSP 2019},
  year = {2019}
}