ICASSP 2025accepted0 citations

A Chinese Expressive Long-dialogue Speech Dataset with Scripts

Jin Li, Tianrui Wang, Meng Ge, Chenrui Cui, Jiahui Zhao, Jianrong Wang, Longbiao Wang, Jianwu Dang

Abstract

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervised data. To address this, we introduce a three-stage data processing pipeline for creating a Chinese expressive long-dialogue speech dataset with scripts (CELSDS). We collect videos from TV series, manually annotate speaker information for each character, apply Optical Character Recognition (OCR) to extract speech content, annotate episode summaries, and use a large language model (LLM) to generate sentence-level scenario descriptions. To our knowledge, this is the first Chinese long-context dialogue dataset that incorporates speaker and content annotations, script-level episode summaries, and sentence-level scenario details. Using this dataset, we develop a baseline model for both speech-to-script and script-to-speech generation tasks. The annotations and data production code are open-sourced at: https://github.com/lijin0120/CELSDS.

BibTeX
@inproceedings{icassp2025_achineseexpressi,
  title = {A Chinese Expressive Long-dialogue Speech Dataset with Scripts},
  author = {Jin Li and Tianrui Wang and Meng Ge and Chenrui Cui and Jiahui Zhao and Jianrong Wang and Longbiao Wang and Jianwu Dang},
  booktitle = {ICASSP 2025},
  year = {2025}
}