ICASSP 2022accepted0 citations

Incorporating End-to-End Framework Into Target-Speaker Voice Activity Detection

Weiqing Wang, Ming Li

Abstract

In this paper, we propose an end-to-end target-speaker voice activity detection (E2E-TS-VAD) method for speaker diarization. First, a ResNet-based network extracts the frame-level speaker embeddings from the acoustic features. Then, the L2-normalized frame-level speaker embeddings are fed to the transformer encoder which produces the initialization of the speaker diarization results. Later, the frame-level speaker embeddings are aggregated to several target-speaker embeddings based on the output from the transformer encoder. Finally, a BiLSTM-based TS-VAD model predicts the refined diarization results. Several aggregation methods are explored, including soft/hard decisions with/without normalization. Results show that E2E-TS-VAD achieves better performance than the original TS-VAD method with the clustering-based initialization.

BibTeX
@inproceedings{icassp2022_incorporatingend,
  title = {Incorporating End-to-End Framework Into Target-Speaker Voice Activity Detection},
  author = {Weiqing Wang and Ming Li},
  booktitle = {ICASSP 2022},
  year = {2022}
}
Incorporating End-to-End Framework Into Target-Speaker Voice Activity Detection · ICASSP 2022