In Search of Optimal Pretraining Strategy for Robust Speaker Recognition
Nikita Khmelev, Stepan Malykh, Alexander Anikin, Anastasia Korenevskaya, Sergey Novoselov, Vladimir Volokhov, Anastasia Zorkina, Vladislav Marchevskiy
Abstract
While demonstrating state-of-the-art results in the microphone channel domain (VoxCeleb protocols), contemporary speaker verification systems are not often tested in challenging acoustic environments such as telephone channel or far-field microphone. This paper compares modern pretraining strategies, proven beneficial for the speaker verification task. It follows wav2vec 2.0, HuBERT, ASR procedures, and aims to identify the most effective, robust approach. We conduct a range of experiments with pretraining on the LibriSpeech corpus and finetuning on the VoxCeleb dataset. The systems are evaluated on multiple protocols with the microphone, telephone, and cross-channel tasks. Our empirical results show that ASR pretraining demonstrates superior in-domain performance but fails to match HuBERT/wav2vec 2.0 in out-of-domain NIST SRE assessment. Adoption of wav2vec 2.0 strategy achieves a 34% average improvement in out-of-domain evaluations compared to the baseline systems. We employ UMAP visualization of models’ embedding space to further understand the reasons for unstable performance in adversarial conditions. We also conclude that while a choice of a pretraining scheme is important, the impact of a speaker verification backend is negligible.
BibTeX
@inproceedings{icassp2025_insearchofoptima,
title = {In Search of Optimal Pretraining Strategy for Robust Speaker Recognition},
author = {Nikita Khmelev and Stepan Malykh and Alexander Anikin and Anastasia Korenevskaya and Sergey Novoselov and Vladimir Volokhov and Anastasia Zorkina and Vladislav Marchevskiy and Galina Lavrentyeva},
booktitle = {ICASSP 2025},
year = {2025}
}