ICASSP 2024accepted0 citations

Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection

Lishi Zuo, Man-Wai Mak, Youzhi Tu

Abstract

Recent studies have demonstrated the effectiveness of speaker disentanglement in mitigating the interference caused by speaker features in speech-based depression detection. However, the inherent entanglement between depression features and speaker features poses challenges to depression detection. In this study, we propose a mutual information-based speaker-invariant depression detector (MI-SIDD) that aims to promote independence between depression and speaker features to facilitate speaker disentanglement. Specifically, we disentangle the speaker features using a vanilla autoencoder with a well-tuned bottleneck layer and minimize the mutual information between depression and speaker features using a conditional mutual information constraint. Experimental results demonstrate the effectiveness of speaker disentanglement and the promotion of independence between depression and speaker features. Our MI-SIDD model achieves competitive performance compared to state-of-the-art methods on the DAIC-WOZ dataset.

BibTeX
@inproceedings{icassp2024_promotingindepen,
  title = {Promoting Independence of Depression and Speaker Features for Speaker Disentanglement in Speech-Based Depression Detection},
  author = {Lishi Zuo and Man-Wai Mak and Youzhi Tu},
  booktitle = {ICASSP 2024},
  year = {2024}
}