Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning
Yuke Lin, Fulin Zhang, Yingying Gao, Shilei Zhang, Ming Li
Abstract
Discrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be largely preserved in the compression process since the reconstructed voice is almost the same in human listening. In this paper, we explore various training strategies and codec types for NAC-based speaker representation learning. Using ECAPA-TDNN as the model backbone, our approach achieves state-of-the-art performance with a 2.08% EER in NAC-based speaker verification scenarios. To better retain speaker information in early, more compressed layers, we introduce mask-layer augmentation and embedding fusion techniques during the training process. Experimental results show the effectiveness of our methods, particularly when inferring with limited codec layers.
BibTeX
@inproceedings{icassp2025_codecasvexplorin,
title = {Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning},
author = {Yuke Lin and Fulin Zhang and Yingying Gao and Shilei Zhang and Ming Li},
booktitle = {ICASSP 2025},
year = {2025}
}