SSE: A Speaking Style Extractor Based on Fine-Grained Contrastive Learning between Speech and Descriptive Text
Zixing Zhang, Yimeng Wu, Zhongren Dong, Wulong Xiang, Shengfan Shen, Björn W. Schuller
Abstract
Effective extraction of paralinguistic features from speech, such as emotion, accent, and age, remains a challenging task in speech processing. Traditional methods typically address each type of paralinguistic information with separate classification or regression tasks—e. g., emotion recognition, accent detection, and age estimation. These approaches can result in fragmented and incomplete descriptions of speech style, failing to capture the full range of paralinguistic attributes in an integrated way. To address these limitations, we propose the Speech Style Extractor (SSE), a novel approach that aims to provide a more comprehensive extraction of speech style features. SSE leverages an optimized fine-grained contrastive learning scheme, enhancing the extraction of diverse paralinguistic features. The optimization approach outperforms the baseline on 11 out of 13 datasets, with an average improvement of 2.1% in terms of the R@1 evaluation metric. We also propose a complete data generation framework using large language models to create 199k text samples across 9 paralinguistic categories, such as emotion, deception, stuttering, and accent, to fill the gap in speaking style descriptions in public datasets. This extensive dataset not only facilitates our research but also serves as a valuable resource for the speech processing community.
BibTeX
@inproceedings{icassp2025_sseaspeakingstyl,
title = {SSE: A Speaking Style Extractor Based on Fine-Grained Contrastive Learning between Speech and Descriptive Text},
author = {Zixing Zhang and Yimeng Wu and Zhongren Dong and Wulong Xiang and Shengfan Shen and Björn W. Schuller},
booktitle = {ICASSP 2025},
year = {2025}
}