2020
Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features
ICASSP 2020accepted
Deep neural network based audio embeddings (d-vectors) have demonstrated superior performance in audio-only speaker diarization compared to traditional acoustic features such as mel-frequency cepstral coefficients (MFCCs) and i-vectors. However, there has been little work on multimodal diarization s…