← Search

Jingdong Li

4 accepted papers

2022

The PCG-AIID System for L3DAS22 Challenge: MIMO and MISO Convolutional Recurrent Network for Multi Channel Speech Enhancement and Speech Recognition

ICASSP 2022accepted

This paper described the PCG-AIID system for L3DAS22 challenge in Task 1: 3D speech enhancement in office reverberant environment. We proposed a two-stage framework to address multi-channel speech denoising and dereverberation. In the first stage, a multiple input and multiple out-put (MIMO) network…

Cited by 0SourceScholar
2022

Uformer: A Unet Based Dilated Complex & Real Dual-Path Conformer Network for Simultaneous Speech Enhancement and Dereverberation

ICASSP 2022accepted

Complex spectrum and magnitude are considered as two major features of speech enhancement and dereverberation. Traditional approaches always treat these two features separately, ignoring their underlying relationship. In this paper, we propose Uformer, a Unet based dilated complex & real dual-path c…

Cited by 0SourceScholar
2021

Densely Connected Multi-Stage Model with Channel Wise Subband Feature for Real-Time Speech Enhancement

ICASSP 2021accepted

Research on single channel speech enhancement (SE) has a long tradition, but two main practical problems still remain unsolved. Firstly, it’s hard to balance between enhancement quality and computational efficiency, and low-latency always brings loss of quality. Secondly, enhancement in specific sce…

Cited by 0SourceScholar
2020

Teacher-Student Training For Robust Tacotron-Based TTS

ICASSP 2020accepted

While neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that resul…

Cited by 0SourceScholar