STE-Mamba: Automated Multimodal Depression Detection through Emotional Analysis and Spatio-Temporal Information Ensemble
Zulong Lin, Yaowei Wang, Yujue Zhou, Fei Du, Yun Yang
Abstract
Automatic Depression Detection (ADD) garners widespread attention due to its convenience and objectivity. While existing research makes significant progress, challenges remain. First, most current ADD methods struggle to balance computational overhead and prediction accuracy. Second, these methods primarily rely on facial images and audio, which are susceptible to external factors, affecting the model’s generalizability. In this study, we propose the Spatiotemporal Ensemble Mamba (STE-Mamba), a framework based entirely on the Mamba architecture for detecting and ensembling spatiotemporal information. This approach reduces computational overhead while effectively capturing long-range spatiotemporal information. Additionally, we extract remote Photoplethysmography (rPPG) and emotion trends (ET) from facial videos, providing two more generalizable physiological modalities for ADD. Experimental results indicate that the inclusion of the ET modality, which only adds two dimensions, improves diagnostic accuracy by approximately 6%. We conduct extensive experiments on five datasets (AVEC2013, AVEC2014, AVEC2017, AVEC2019, CMDep), and the results demonstrate that STE-Mamba is highly competitive in terms of both effectiveness and generalizability. The self-built CMDep can be requested via the following link.
BibTeX
@inproceedings{icassp2025_stemambaautomate,
title = {STE-Mamba: Automated Multimodal Depression Detection through Emotional Analysis and Spatio-Temporal Information Ensemble},
author = {Zulong Lin and Yaowei Wang and Yujue Zhou and Fei Du and Yun Yang},
booktitle = {ICASSP 2025},
year = {2025}
}