Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information
Yongtai Ji, Guangxing Li, Hao Huang, Yanbing Li, Wushour Silamu
Abstract
Speaker verification performance significantly degrades when there exists a language mismatch between training and evaluation. Domain Adversarial Training (DAT) has shown to be effective in mitigating this gap by incorporating adversarial training with domain information (language id). Inspired by recent research of DAT, we propose to decompose and locate features sensitive to speaker identity, so that domain adaptation can better serve the speaker classification task. Additionally, we further promote DAT by replacing alignment based on language identification with alignment of fine-grained phonetic information, where pre-trained speech recognition models are utilized to provide frame-level phonetic labels. The speaker verification model is trained on VoxCeleb, while CnCeleb is used for adversarial training and evaluation. Results show the methods effectively mitigate the performance degradation caused by language mismatch.
BibTeX
@inproceedings{icassp2025_improvedcrosslin,
title = {Improved Cross-Lingual Speaker Verification Using Speaker Sensitive Feature Guidance and Fine-grained Phonetic Information},
author = {Yongtai Ji and Guangxing Li and Hao Huang and Yanbing Li and Wushour Silamu},
booktitle = {ICASSP 2025},
year = {2025}
}