ICASSP 2018accepted0 citations

A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions

Lei Sun, Jun Du, Tian Gao, Yu-Ding Lu, Yu Tsao, Chin-Hui Lee, Neville Ryant

Abstract

In this study, we investigate on the effects of deep learning based speech enhancement as a preprocessor to speaker diarization in quite challenging realistic environments involving the background noises, reverberations and overlapping speech. To improve the generalization capability, the advanced long short-term memory (LSTM) architecture with the novel design of hidden layers via densely connected progressive learning and output layer via multiple-target learning is proposed for preprocessing. We build the deep model using synthesized training data pairs generated from WSJO reading-style speech and more than 100 noise types. Surprisingly, this proposed preprocessor demonstrates a strong generalization capability to speaker di-arization with the realistic noisy speech in highly mismatched conditions, in terms of the speaking style, interferences, and the interaction between them. Tested on three challenging tasks, namely AMI, ADOS, and SeedLings, the state-of-the-art diarization system with the novel LSTM-based speech preprocessor can yield consistent and significant reductions of diarization error rate (DER) over the systems using unprocessed noisy speech and traditional enhancement methods.

BibTeX
@inproceedings{icassp2018_anovellstmbaseds,
  title = {A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions},
  author = {Lei Sun and Jun Du and Tian Gao and Yu-Ding Lu and Yu Tsao and Chin-Hui Lee and Neville Ryant},
  booktitle = {ICASSP 2018},
  year = {2018}
}
A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions · ICASSP 2018