Multi-speaker conversations, cross-talk, and diarization for speaker recognition
Abstract
I-vector training and extraction assume that a speech file is spoken by a single speaker. This work considers the effects of violating that assumption with the presence of cross-talk or multi-speaker conversations. First, it is demonstrated that these problematic speech files can be detected using the i-vector representation itself. The impact of these violations of the single-speaker assumption are then explored along with strategies to mitigate it. It is shown that, even in predominantly clean data, the removal of cross-talk can provide modest gains, but that T matrix and PLDA training are largely robust to these types of noise. It is also shown that detection in front of diarization is a reasonable strategy in the presence of data with an unknown proportion of multi-speaker conversations. Finally, in the course of this work, evidence is found that cross-talk detection and multi-speaker detection may in fact be different tasks that require separately trained detectors.
BibTeX
@inproceedings{icassp2017_multispeakerconv,
title = {Multi-speaker conversations, cross-talk, and diarization for speaker recognition},
author = {Gregory Sell and Alan McCree},
booktitle = {ICASSP 2017},
year = {2017}
}