2020
Self-Supervised Learning for Audio-Visual Speaker Diarization
ICASSP 2020accepted
Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propose a self-supervised audio-video synchronization learning method to address the…