DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming
Shlomo E. Chazan, Jacob Goldberger, Sharon Gannot
Abstract
In this paper, we present a new control mechanism for LCMV beamforming. Application of the LCMV beamformer to speaker separation tasks requires accurate estimates of its building blocks, e.g. the noise spatial cross-power spectral density (cPSD) matrix and the relative transfer function (RTF) of all sources of interest. An accurate classification of the input frames to various speaker activity patterns can facilitate such an estimation procedure. We propose a DNN-based concurrent speakers detector (CSD) to classify the noisy frames. The CSD, trained in a supervised manner using a DNN, classifies noisy frames into three classes: 1) all speakers are inactive - used for estimating the noise spatial cPSD matrix; 2) a single speaker is active - used for estimating the RTF of the active speaker; and 3) more than one speaker is active - discarded for estimation purposes. Finally, using the estimated blocks, the LCMV beamformer is constructed and applied for extracting the desired speaker from a noisy mixture of speakers.
BibTeX
@inproceedings{icassp2018_dnnbasedconcurre,
title = {DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming},
author = {Shlomo E. Chazan and Jacob Goldberger and Sharon Gannot},
booktitle = {ICASSP 2018},
year = {2018}
}