Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison
Min Hyun Han, Sung Hwan Mun, Minchan Kim, Myeonghun Jeong, Sunghwan Ahn, Nam Soo Kim
Abstract
Deep speaker embedding systems are usually trained with classification-based or end-to-end learning objectives. Popular end-to-end approaches utilize deep metric learning, which can be viewed as a few-shot classification objective. In this paper, we investigate the limit of conventional learning objectives in speaker verification and propose a new learning objective designed from the perspective of similarity scores. The proposed method trains a network by score comparison unbound from the classification, which is more suitable for verification tasks. Experiments conducted with popular speaker embedding networks demonstrate the improvements on the VoxCeleb dataset using the proposed loss.
BibTeX
@inproceedings{icassp2023_improvinglearnin,
title = {Improving Learning Objectives for Speaker Verification from the Perspective of Score Comparison},
author = {Min Hyun Han and Sung Hwan Mun and Minchan Kim and Myeonghun Jeong and Sunghwan Ahn and Nam Soo Kim},
booktitle = {ICASSP 2023},
year = {2023}
}