Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment
Abinay Reddy Naini, Donita Robinson, Elizabeth Richerson, Carlos Busso
Abstract
This work addresses the challenge of building speech emotion recognition models that generalize effectively across different domains, particularly when only limited target domain data is available with or without emotional label information. Traditional models often struggle with cross-domain performance due to the variability in emotional expressions and the lack of alignment between the training and target domains. We propose a novel approach that prioritizes aligning the emotional label distribution of the training data with that of the target domain by undersampling the source domain. Even though we intentionally reduce the size of the training set from the source domain, the emotional content alignment leads to clear performance improvements, outperforming models trained with the complete training set. This strategy highlights the importance of aligning emotional attributes during training, helping to create robust emotion recognition models across diverse applications. Our findings also reveal that performance significantly improves when even a small amount of labeled target domain data is available, allowing for a more accurate assessment of the emotional distribution in the target domain.
BibTeX
@inproceedings{icassp2025_domainspecificad,
title = {Domain-Specific Adaptation in Speech Emotion Recognition Using Emotional Distribution Alignment},
author = {Abinay Reddy Naini and Donita Robinson and Elizabeth Richerson and Carlos Busso},
booktitle = {ICASSP 2025},
year = {2025}
}