SSC: 106 bit/s Ultra-low Bitrate Semantic Speech Coding
Renjie Jia, Zhiqiang He, Kai Niu, Zixuan Xiao, Yonghui Liu, Jianbing Liu
Abstract
Currently, balancing low bitrate coding with speech quality is a highly debated topic in the research community. At very low bitrates, existing methods often fail to maintain speech naturalness, intelligibility, and personalization. To address this issue, we introduce an innovative ultra-low bitrate semantic speech coding approach, termed Semantic Speech Coding (SSC). Specifically, the multi-level feature extraction and compression mechanism sequentially extracts and compresses speech features at different levels, ensuring speech quality at ultra-low bit rates. Using a semantic vector quantization codec to fuse spectral and pitch features to extract essential semantic information, achieving more efficient compression while enhancing intelligibility and naturalness. The low-data-overhead speaker feature encoder captures time-invariant speaker characteristics, enabling personalized speech synthesis without additional data overhead, ensuring the synthesized speech retains personalization and naturalness. The diffusion loss mechanism employs a conditional diffusion model to progressively restore details, mitigating the detail loss typically seen in conventional codecs, further enhancing the naturalness and realism of the synthesized speech. We achieved significant improvements in speech quality at an ultra-low bitrate of 106 bps, which approaches the theoretical upper limit of information rate.
BibTeX
@inproceedings{icassp2025_ssc106bitsultral,
title = {SSC: 106 bit/s Ultra-low Bitrate Semantic Speech Coding},
author = {Renjie Jia and Zhiqiang He and Kai Niu and Zixuan Xiao and Yonghui Liu and Jianbing Liu},
booktitle = {ICASSP 2025},
year = {2025}
}