ICASSP 2025accepted0 citations

A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification

Shengyu Peng, Wu Guo, Jie Zhang, Zuoliang Li, Yu Guan, Bin Gu, Yang Ai

Abstract

In this paper, a multi-scale feature fusion paradigm is proposed to fully exploit the power of the pre-trained models for text-independent speaker verification. It contains a front-end feature extractor and an enhanced ECAPA-TDNN backend in a cascade manner. The feature extractor incorporates local representations of the CNN layers as well as the global clues of the Transformer layers of the pre-trained models, which are combined to construct the multi-scale discriminative features. The outputs of the feature extractor are then fed into the back-end model (tailored from ECAPA-TDNN) to obtain the final speaker embedding. Results on VoxCeleb datasets validate the superiority of the proposed method with equal error rates of 0.633% and 0.457% on the official trials of Vox1-O using the base and large pre-trained models, respectively.

BibTeX
@inproceedings{icassp2025_astudyofmultisca,
  title = {A Study of Multi-Scale Feature Learning From Pre-Trained Models on Speaker Verification},
  author = {Shengyu Peng and Wu Guo and Jie Zhang and Zuoliang Li and Yu Guan and Bin Gu and Yang Ai},
  booktitle = {ICASSP 2025},
  year = {2025}
}