← Search

Shidong Shang

8 accepted papers

2025

AVS3P10 Standard for Real-time Speech Coding

ICASSP 2025accepted

As the tenth part of the third-generation AVS standard series for real-time speech coding, AVS3P10 is the recent standard completed in the Audio Video Coding Standards Workgroup of China (AVS). Combining the state-of-the-art deep generative networks and signal processing methods, AVS3P10 targets def…

Cited by 0SourceScholar
2025

Full-text Error Correction for Chinese Speech Recognition with Large Language Model

ICASSP 2025accepted

Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper i…

Cited by 0SourceScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2023

Inter-Subnet: Speech Enhancement with Subband Interaction

ICASSP 2023accepted

Subband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral…

Cited by 0SourceScholar
2023

Speech Enhancement with Intelligent Neural Homomorphic Synthesis

ICASSP 2023accepted

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral ana…

Cited by 0SourceScholar
2023

TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge

ICASSP 2023accepted

This paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to…

Cited by 0SourceScholar
2022

Internet Streaming Audio Based Speech Reception Threshold Measurement in Cochlear Implant Users

ICASSP 2022accepted

Traditional face-to-face subjective listening test has become a challenge due to the COVID-19 pandemic. We developed a remote assessment system with Tencent Meeting, a video conferencing application, to address this issue. This paper presents our work on evaluating the reliability of the remote asse…

Cited by 0SourceScholar
2022

TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge

ICASSP 2022accepted

This paper describes Tencent Ethereal Audio Lab – Northwestern Polytechnical University personalized speech enhancement (TEA-PSE) system submitted to track 2 of the ICASSP 2022 Deep Noise Suppression (DNS) challenge. Our system specifically combines the dual-stage network which is a superior real-ti…

Cited by 56SourceScholar