CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification
Heng Li, Yukun Qian, Yun Lu, Mingjiang Wang
Abstract
Sleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the acoustic characteristics of snoring vary significantly among individuals. To address this issue, we introduced a text-audio multimodal model that leverages patient’s metadata to provide valuable supplementary information for SASC task. Specifically, we utilized text descriptions derived from metadata and snoring sounds to fine-tune a pretrained text-audio multimodal model. The metadata includes patient’s physical indicators such as gender, age, BMI, neck circumference, and blood pressure. We constructed a snoring dataset that included four sleep apnea severity levels. On this dataset, our method achieved a classification F-score of 74.34%. We conducted a series of ablation experiments to validate the effectiveness of improving SASC performance by leveraging both metadata-based text and snoring sounds. Additionally, we discussed the model’s performance in scenarios where parts of the metadata are unavailable, a situation that may occur in real-world applications.
BibTeX
@inproceedings{icassp2025_csmtcombiningsno,
title = {CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification},
author = {Heng Li and Yukun Qian and Yun Lu and Mingjiang Wang},
booktitle = {ICASSP 2025},
year = {2025}
}