Using Emotionally Rich Speech Segments for Depression Prediction
Abstract
The automatic assessment of depression through human voice has gained increasing interest due to its cost-effectiveness and non-invasiveness. This paper employs acoustic embeddings from emotionally rich speech segments for depression prediction, given the strong connection between depression and emotion expression. We leverage a large pre-trained model to predict depression. The public dimensional emotion model (PDEM) used in this study is fine-tuned for recognizing arousal, valence and dominance. We use PDEM for both extracting embeddings and selecting emotionally rich speech segments based on its arousal, valence, and dominance predictions. We advance the state-of-the-art performance on both the Androids corpus (Interview task) for depression detection, following the predetermined protocol, and the E-DAIC corpus for acoustic-based depression severity prediction, adhering to the 2019 Audio Visual Emotion Challenge (AVEC) protocol. The analysis demonstrates that emotionally rich speech segments contain more depression-related cues compared to emotion-neutral segments.
BibTeX
@inproceedings{icassp2025_usingemotionally,
title = {Using Emotionally Rich Speech Segments for Depression Prediction},
author = {Jiawei Yu and Heysem Kaya},
booktitle = {ICASSP 2025},
year = {2025}
}