ICASSP 2018accepted0 citations

Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion Recognition

Saurabh Sahu, Rahul Gupta, Ganesh Sivaraman, Carol Y. Espy-Wilson

Abstract

Training discriminative classifiers involves learning a conditional distribution p(y <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sup> |x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> ), given a set of feature vectors x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> and the corresponding labels y <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> , i=1...N. For a classifier to be generalizable and not overfit to training data, the resulting conditional distribution p(y <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> |x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> ) is desired to be smoothly varying over the inputs x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> . Adversarial training procedures enforce this smoothness using manifold regularization techniques. Manifold regularization makes the model's output distribution more robust to local perturbation added to a datapoint x <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">i</sub> . In this paper, we experiment with the application of adversarial training procedures to increase the accuracy of a deep neural network based emotion recognition system using speech cues. Specifically, we investigate two training procedures: (i) adversarial training where we determine the adversarial direction based on the given labels for the training data and, (ii) virtual adversarial training where we determine the adversarial direction based only on the output distribution of the training data. We demonstrate the efficacy of adversarial training procedures by performing a k-fold cross validation experiment on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) and a cross-corpus performance analysis on three separate corpora. Results show improvement over a purely supervised approach, as well as better generalization capability to cross-corpus settings.

BibTeX
@inproceedings{icassp2018_smoothingmodelpr,
  title = {Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion Recognition},
  author = {Saurabh Sahu and Rahul Gupta and Ganesh Sivaraman and Carol Y. Espy-Wilson},
  booktitle = {ICASSP 2018},
  year = {2018}
}
Smoothing Model Predictions Using Adversarial Training Procedures for Speech Based Emotion Recognition · ICASSP 2018