A Novel Audio-Visual Multimodal Semi-Supervised Model Based on Graph Neural Networks for Depression Detection
Yaqin Li, Chenjian Sun, Yihong Dong
Abstract
There is a significant correlation between depression, verbal behavior, and facial expressions. By analyzing patients’ audio and facial visuals, depression assessments can be conducted. However, existing work is predominantly based on single modalities. Additionally, acquiring a sufficient amount of labeled data in clinical settings is challenging and costly. To leverage multimodal audio-visual data while addressing the issue of lacking trainable labeled data, we propose an audiovisual multimodal semi-supervised depression detection model based on Graph Neural Networks (AVS-GNN). This model first extracts dual-modality temporal information from audio features and facial visual features of patients and obtains modality-specific high-level embedding representations through graph representation learning. Subsequently, it utilizes graph-based contrastive unsupervised learning to capture consistency information between pairs of unlabeled samples across different modalities and to facilitate cross-modal interactions. We specifically designed a hybrid weighted pseudo-labeling strategy to assign high-confidence pseudo-labels to unlabeled data and further retraining the model. Experiments on two depression datasets show that this model outperforms baseline methods across all evaluation metrics.
BibTeX
@inproceedings{icassp2025_anovelaudiovisua,
title = {A Novel Audio-Visual Multimodal Semi-Supervised Model Based on Graph Neural Networks for Depression Detection},
author = {Yaqin Li and Chenjian Sun and Yihong Dong},
booktitle = {ICASSP 2025},
year = {2025}
}