Within-Sample Variability-Invariant Loss for Robust Speaker Recognition Under Noisy Environments
Despite the significant improvements in speaker recognition enabled by deep neural networks, unsatisfactory performance persists under noisy environments. In this paper, we train the speaker embedding network to learn the "clean" embedding of the noisy utterance. Specifically, the network is trained…