ICASSP 2025accepted0 citations

AVS3P10 Standard for Real-time Speech Coding

Wei Xiao, Weibei Dou, Wenlong Wang, Gaoxiong Yi, Jingxin Li, Shidong Shang

Abstract

As the tenth part of the third-generation AVS standard series for real-time speech coding, AVS3P10 is the recent standard completed in the Audio Video Coding Standards Workgroup of China (AVS). Combining the state-of-the-art deep generative networks and signal processing methods, AVS3P10 targets defining new generation neural speech codecs with high quality at low bitrates, enabling excellent experiences even when the bitrate is at 5.9 kbps with excellent error resilience. Moreover, it provides wideband and super wideband coding modes, and it supports the extension of stereo coding. Both subjective listening test and objective measurement prove the merit of AVS3P10. Especially, a lightweight model with only 880k parameters is incorporated to maintain the practicality of AVS3P10 in computational efficiency. Conclusively, AVS3P10 demonstrates the maturity of neural speech coding with broad application perspectives in real-time communication.

BibTeX
@inproceedings{icassp2025_avs3p10standardf,
  title = {AVS3P10 Standard for Real-time Speech Coding},
  author = {Wei Xiao and Weibei Dou and Wenlong Wang and Gaoxiong Yi and Jingxin Li and Shidong Shang},
  booktitle = {ICASSP 2025},
  year = {2025}
}