Robust Self-Supervised Speaker Representation Learning Via Instance Mix Regularization
Woo Hyun Kang, Jahangir Alam, Abderrahim Fathan
Abstract
Over the recent years, various self-supervised contrastive embedding learning methods for deep speaker verification were proposed. The performance of the self-supervised contrastive learning framework highly depends on the data augmentation technique, but due to the sensitive nature of speaker information within the speech signal, most speaker embedding training relies on simple augmentations such as additive noise or simulated reverberation. Thus while the conventional self-supervised speaker embedding systems can yield minimum within-utterance variability, the capability to generalize to out-of-set utterance is limited. In order to alleviate this problem, we propose a novel self-supervised learning framework for speaker verification which combines the angular prototypical loss and the instance mix (i-mix) regularization. The proposed method was evaluated on the VoxCeleb1 dataset and showed noticeable improvement over the standard self-supervised embedding method.
BibTeX
@inproceedings{icassp2022_robustselfsuperv,
title = {Robust Self-Supervised Speaker Representation Learning Via Instance Mix Regularization},
author = {Woo Hyun Kang and Jahangir Alam and Abderrahim Fathan},
booktitle = {ICASSP 2022},
year = {2022}
}