SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification
Yishuang Li, Hukai Huang, Zhicong Chen, Wenhao Guan, Jiayan Lin, Lin Li, Qingyang Hong
Abstract
Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly models speaker information — Speaker Related HuBERT, abbreviated as SR-HuBERT. This framework aims to further explore speaker-related information inherent in speech universal representations. The proposed SR-HuBERT utilizes an unsupervised clustering algorithm based on graph structures to generate speaker pseudo-labels and promotes the learning of segment-level speaker-related representations through a multi-task pre-training framework. Experimental results on VoxCeleb1 test set demonstrate the effectiveness of the proposed SR-HuBERT. Even in the scenarios of limited fine-tuning data, SR-HuBERT outperforms the other existing PTMs on SV tasks. Additionally, SR-HuBERT also performs well on speaker-related tasks of SUPERB benchmark.
BibTeX
@inproceedings{icassp2024_srhubertaneffici,
title = {SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification},
author = {Yishuang Li and Hukai Huang and Zhicong Chen and Wenhao Guan and Jiayan Lin and Lin Li and Qingyang Hong},
booktitle = {ICASSP 2024},
year = {2024}
}