Learning Strategy with Barlow Twins Objective for Emotion-Robust Speaker Verification System
Dmitrii V. Mikhailovskii, Aki Kunikoshi, David A. van Leeuwen, Jaebok Kim
Abstract
A recent trend of the automatic speaker verification (ASV) systems is to use speaker representations extracted from deep learning-based speaker encoders. Those representations are robust against linguistic variation but sensitive to emotional fluctuation that potentially degrades the performance of ASV systems. To tackle this issue, we applied a correlation-based objective function called Barlow Twins objective to speaker representation learning for the ASV task with expressive speech. This helps the learned representations from the same speaker become similar despite their emotional state. Our experiment showed that our representation learning using Barlow Twins objective improves the standard deviation of the EERs between different emotions from 2.43 to 1.52 and the total EER from 10.77% to 6.47% on the CREMA-D dataset. The study also evaluates the system’s generalization performance across out-of-domain datasets, demonstrating improved standard deviation of the EERs and total EER of the baseline system in all datasets while observing the performance drop on synthetic emotional data where emotional mismatch occurs.
BibTeX
@inproceedings{icassp2025_learningstrategy,
title = {Learning Strategy with Barlow Twins Objective for Emotion-Robust Speaker Verification System},
author = {Dmitrii V. Mikhailovskii and Aki Kunikoshi and David A. van Leeuwen and Jaebok Kim},
booktitle = {ICASSP 2025},
year = {2025}
}