ICASSP 2023accepted0 citations

Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition

Li-Wei Chen, Alexander Rudnicky

Abstract

While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla fine-tuning (V-FT) and task adaptive pretraining (TAPT) are first presented. We show that V-FT is able to outperform state-of-the-art models on the IEMOCAP dataset. TAPT, an existing NLP fine-tuning strategy, further improves the performance on SER. We also introduce a novel fine-tuning method termed P-TAPT, which modifies the TAPT objective to learn contextualized emotion representations. Experiments show that P-TAPT performs better than TAPT, especially under low-resource settings. Compared to prior works in this literature, our top-line system achieved a 7.4% absolute improvement in unweighted accuracy (UA) over the state-of-the-art performance on IEMOCAP. Our code is publicly available. <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup>

BibTeX
@inproceedings{icassp2023_exploringwav2vec,
  title = {Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition},
  author = {Li-Wei Chen and Alexander Rudnicky},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition · ICASSP 2023