ICASSP 2015accepted0 citations

Softsad: Integrated frame-based speech confidence for speaker recognition

Mitchell McLaren, Martin Graciarena, Yun Lei

Abstract

In this paper we propose softSAD: the direct integration of speech posteriors into a speaker recognition system as an alternative to using speech activity detection (SAD). Motivated by the need to use audio from short recordings more efficiently, softSAD removes the need to discard audio using speech/non-speech decisions based on a threshold as done with SAD. Instead, softSAD explicitly integrates into the Baum-Welch statistics a speech posterior for each frame. We compare softSAD and SAD in mismatched conditions by evaluating a system developed for the National Institute for Standards and Technology (NIST) 2012 speaker recognition evaluation (SRE) on the short test conditions of the channel-degraded Robust Automatic Transcription of Speech (RATS) speaker identification task (and vice versa). We demonstrate that softSAD provides benefit over SAD for short test audio in mismatched conditions.

BibTeX
@inproceedings{icassp2015_softsadintegrate,
  title = {Softsad: Integrated frame-based speech confidence for speaker recognition},
  author = {Mitchell McLaren and Martin Graciarena and Yun Lei},
  booktitle = {ICASSP 2015},
  year = {2015}
}
Softsad: Integrated frame-based speech confidence for speaker recognition · ICASSP 2015