Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features
Wonjune Kang, Brandon C. Roy, Wesley Chow
Abstract
Deep neural network based audio embeddings (d-vectors) have demonstrated superior performance in audio-only speaker diarization compared to traditional acoustic features such as mel-frequency cepstral coefficients (MFCCs) and i-vectors. However, there has been little work on multimodal diarization systems that combine dvectors with additional sources of information. In this paper, we present a novel approach to multimodal speaker diarization that combines d-vectors with spatial information derived from performing beamforming given a multi-channel microphone array. Our system performs spectral clustering on a combination of speaker embeddings and spatial features that are computed using the Steered-Response Power Phase Transform (SRP-PHAT) algorithm. We evaluate our system on the AMI Meeting Corpus and an internal dataset of real-world conversations. By using both acoustic and spatial features for diarization, we achieve significant improvements over a dvector only baseline and show potential to achieve comparable results with other state-of-the-art multimodal diarization systems.
BibTeX
@inproceedings{icassp2020_multimodalspeake,
title = {Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features},
author = {Wonjune Kang and Brandon C. Roy and Wesley Chow},
booktitle = {ICASSP 2020},
year = {2020}
}