2024
Large Scale Self-Supervised Pretraining for Active Speaker Detection
ICASSP 2024accepted
In this work we investigate the impact of a large-scale self-supervised pretraining strategy for active speaker detection (ASD) on an unlabeled dataset consisting of over 125k hours of YouTube videos. When compared to a baseline trained from scratch on much smaller in-domain labeled datasets we show…