Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel Weighting
Ke Liu, Dekui Wang, Dongya Wu, Jun Feng
Abstract
Multi-view Speech Emotion Recognition (SER) based on the pre-trained model has achieved success in speaker-independent scenarios. However, the existing SER methods rely on excessive feature views and have complicated feature fusion strategies. In this paper, we propose a novel method to learn effective emotion-related information from two feature views. First, we present a Discriminative Channel Weighting (DCW) module to weight the channel dimension of the features produced by a set of multi-scale convolution layers. This module allows for discriminative weighting of complex channel dimensions. Second, a concise Two-stream Pooling Attention (TsPA) strategy is proposed to generate two groups of fusion features based on different channel-level embeddings with different emphasis. Finally, the SER task is completed by three consecutive fully connected layers. The effectiveness of the proposed method has been demonstrated on two speaker-independent validation strategies, outperforming other state-of-the-art approaches.
BibTeX
@inproceedings{icassp2023_speechemotionrec,
title = {Speech Emotion Recognition Via Two-Stream Pooling Attention With Discriminative Channel Weighting},
author = {Ke Liu and Dekui Wang and Dongya Wu and Jun Feng},
booktitle = {ICASSP 2023},
year = {2023}
}