Multi-Source Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition
Shenjie Jiang, Peng Song, Shaokai Li, Run Wang, Wenming Zheng
Abstract
As an important research direction in the field of speech signal processing, cross-domain speech emotion recognition (SER) has attracted extensive attention. In practice, it is challenging to collect enough labeled samples from single source domain to train robust classifiers. To this end, this paper presents a novel method named multi-source unsupervised transfer components learning (MUTCL) for cross-domain SER. In MUTCL, we first adopt a PCA-like strategy and apply it to multi-source domains, aiming to preserve both intra-domain individuality and inter-domain commonality principal components within each domain. Simultaneously, a simple alignment strategy is developed to guide cross-domain samples to have similar structures, thus preserving more transfer components. Moreover, an adaptive weight strategy is utilized to determine the contribution of each source domain. We conduct experiments on five benchmark datasets, and the results show that MUTCL achieves excellent performance compared with some state-of-the-art methods.
BibTeX
@inproceedings{icassp2024_multisourceunsup,
title = {Multi-Source Unsupervised Transfer Components Learning for Cross-Domain Speech Emotion Recognition},
author = {Shenjie Jiang and Peng Song and Shaokai Li and Run Wang and Wenming Zheng},
booktitle = {ICASSP 2024},
year = {2024}
}