ICASSP 2022accepted0 citations

Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition

Tiantian Feng, Hanieh Hashemi, Murali Annavaram, Shrikanth S. Narayanan

Abstract

Speech Emotion Recognition (SER) techniques have gained considerable interest in many applications including smart virtual assistants and health state tracking. SER systems often acquire and transmit speech data collected at the client-side to remote cloud platforms for inference and decision making. However, speech data carries rich information not only about emotions conveyed in vocal expressions, but also other sensitive demographic traits, such as gender, age, and language background. It is desirable to select only features that are necessary for the emotion classification while protecting sensitive features. However, there are some features that are necessary for emotion classification. These features may also reveal other demographic traits. In this work, we propose a method to improve inference privacy for sensitive features by injecting noise into the input speech data, but without degrading the SER system performance. The approach combines a noise representation learning architecture, called Cloak [1], with adversarial training to keep relevant information inside the data for emotion classification while removing information that would enable inferring sensitive demographic attributes. Experimental results show that our method can effectively prevent inference of sensitive demographic information, and that the improved privacy comes at a cost of only a minor utility loss for the emotion classification.

BibTeX
@inproceedings{icassp2022_enhancingprivacy,
  title = {Enhancing Privacy Through Domain Adaptive Noise Injection For Speech Emotion Recognition},
  author = {Tiantian Feng and Hanieh Hashemi and Murali Annavaram and Shrikanth S. Narayanan},
  booktitle = {ICASSP 2022},
  year = {2022}
}