ICASSP 2023accepted0 citations

The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge

Ming Cheng, Haoxu Wang, Ziteng Wang, Qiang Fu, Ming Li

Abstract

This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio-visual signals. The final system achieves a diarization error rate (DER) of 8.82% on the evaluation set of the competition database, which ranks 1st in the speaker diarization track of the MISP 2022, ICASSP Signal Processing Grand Challenge.

BibTeX
@inproceedings{icassp2023_thewhualibabaaud,
  title = {The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge},
  author = {Ming Cheng and Haoxu Wang and Ziteng Wang and Qiang Fu and Ming Li},
  booktitle = {ICASSP 2023},
  year = {2023}
}
The WHU-Alibaba Audio-Visual Speaker Diarization System for the MISP 2022 Challenge · ICASSP 2023