Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations
Zuoliang Li, Yang Ai, Jie Zhang, Shengyu Peng, Yu Guan, Bin Gu, Wu Guo
Abstract
In this paper, we propose a noise-invariant speaker representation learning (SRL) approach by aligning noisy-clean speech pairs at both the feature and embedding levels for model training. Specifically, we first construct noisy-clean pairs using data augmentation during training. The noisy features are then processed by a Conformer-based enhancement module. The feature-level alignment is achieved by minimizing the mean squared error between the enhanced and original clean data. At the embedding level, we introduce a supervised contrastive learning loss with noise-adaptive margin to simultaneously enhance the intra-speaker compactness and the inter-speaker separability and better adapt different noise levels, in combination with the Barlow Twins self-supervised loss to align the noisy-clean data pairs and reduce noise redundancy in the embedding space. Finally, these loss components are integrated with conventional classification loss to train the SRL network. Experimental results on various VoxCeleb1 test sets synthesized with noise sources demonstrate the effectiveness of the proposed method.
BibTeX
@inproceedings{icassp2025_aligningnoisycle,
title = {Aligning Noisy-Clean Speech Pairs at Feature and Embedding Levels for Learning Noise-Invariant Speaker Representations},
author = {Zuoliang Li and Yang Ai and Jie Zhang and Shengyu Peng and Yu Guan and Bin Gu and Wu Guo},
booktitle = {ICASSP 2025},
year = {2025}
}