← Search

Deliang Wang

60 accepted papers

2025

Elevating Robust ASR By Decoupling Multi-Channel Speaker Separation and Speech Recognition

ICASSP 2025accepted

Despite the tremendous success of automatic speech recognition (ASR) with the introduction of deep learning, its performance is still unsatisfactory in many real-world multi-talker scenarios. Speaker separation excels in separating individual talkers but, as a frontend, it introduces processing arti…

Cited by 0SourceScholar
2025

Extended LSTMs for Knowledge Tracing: Peeking Inside the Black Box (Student Abstract)

AAAI 2025technical

This paper proposes extended Long Short-Term Memory (LSTM) networks for the knowledge tracing task and employs explainable AI methods to address interpretability issues. Specifically, we developed an extended LSTM-based model to automatically diagnose students' knowledge states. We then leveraged th…

Cited by 0SourcePDFScholar
2025

Robust Frame-level Speaker Localization in Reverberant and Noisy Environments by Exploiting Phase Difference Losses

ICASSP 2025accepted

This paper investigates robust speaker localization at the frame level on the basis of complex spectral mapping, which is capable of learning both the magnitude and phase of the target signal. Unlike prevailing deep learning methods for speaker localization, we perform MIMO (multi-input multi-output…

Cited by 0SourceScholar
2024

Audiovisual Speaker Separation with Full- and Sub-Band Modeling in the Time-Frequency Domain

ICASSP 2024accepted

We introduce a new deep learning model for talker-independent audiovisual speaker separation in noisy conditions in the time-frequency domain. The inputs to the model include noisy multi-talker mixtures and the corresponding cropped face images. Our approach incorporates cross-attention audiovisual…

Cited by 0SourceScholar
2024

Leveraging Sound Localization to Improve Continuous Speaker Separation

ICASSP 2024accepted

Continuous speaker separation aims to separate overlapping speakers in real-world environments like meetings, but it often falls short in isolating speech segments of a single speaker. This leads to split signals that adversely affect downstream applications such as automatic speech recognition and…

Cited by 9SourceScholar
2023

DATA2VEC-SG: Improving Self-Supervised Learning Representations for Speech Generation Tasks

ICASSP 2023accepted

Self-supervised learning has been successfully applied to various speech recognition and understanding tasks. However, for generative tasks such as speech enhancement and speech separation, most self-supervised speech representations did not show substantial improvements. To deal with this problem,…

Cited by 0SourceScholar
2023

Multi-Resolution Location-Based Training for Multi-Channel Continuous Speech Separation

ICASSP 2023accepted

The performance of automatic speech recognition (ASR) systems severely degrades when multi-talker speech overlap occurs. In meeting environments, speech separation is typically performed to improve the robustness of ASR systems. Recently, location-based training (LBT) was proposed as a new training…

Cited by 8SourceScholar
2022

Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex Domain

ICASSP 2022accepted

Bone-conduction (BC) microphones capture speech signals by converting the vibrations of the human skull into electrical signals. BC sensors are insensitive to acoustic noise, but limited in bandwidth. On the other hand, conventional or air-conduction (AC) microphones are capable of capturing full-ba…

Cited by 0SourceScholar
2022

Improving Noise Robustness of Contrastive Speech Representation Learning with Speech Reconstruction

ICASSP 2022accepted

Noise robustness is essential for deploying automatic speech recognition (ASR) systems in real-world environments. One way to reduce the effect of noise interference is to employ a preprocessing module that conducts speech enhancement, and then feed the enhanced speech to an ASR backend. In this wor…

Cited by 0SourceScholar
2022

Location-Based Training for Multi-Channel Talker-Independent Speaker Separation

ICASSP 2022accepted

Permutation-invariant training (PIT) is a dominant approach for addressing the permutation ambiguity problem in talker-independent speaker separation. Leveraging spatial information afforded by microphone arrays, we propose a new training approach to resolving permutation ambiguities for multi-chann…

Cited by 0SourceScholar
2022

Multichannel Speech Enhancement Without Beamforming

ICASSP 2022accepted

Deep neural networks are often coupled with traditional spatial filters, such as MVDR beamformers for effectively exploiting spatial information. Even though single-stage end-to-end supervised models can obtain impressive enhancement, combining them with a traditional beamformer and a DNN-based post…

Cited by 24SourceScholar
2022

Summary on the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Grand Challenge

ICASSP 2022accepted

The ICASSP 2022 Multi-channel Multi-party Meeting Transcription Grand Challenge (M2MeT) focuses on one of the most valuable and the most challenging scenarios of speech technologies. The M2MeT challenge has particularly set up two tracks, speaker diarization (track 1) and multi-speaker automatic spe…

Cited by 0SourceScholar
2022

TPARN: Triple-Path Attentive Recurrent Network for Time-Domain Multichannel Speech Enhancement

ICASSP 2022accepted

In this work, we propose a new model called triple-path attentive recurrent network (TPARN) for multichannel speech enhancement in the time domain. TPARN extends a single-channel dual-path network to a multichannel network by adding a third path along the spatial dimension. First, TPARN processes sp…

Cited by 0SourceScholar
2021

Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral Mapping

ICASSP 2021accepted

Speech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance.…

Cited by 4SourceScholar
2021

Time-Domain Loss Modulation Based on Overlap Ratio for Monaural Conversational Speaker Separation

ICASSP 2021accepted

Existing speaker separation methods deliver excellent performance on fully overlapped signal mixtures. To apply these methods in daily conversations that include occasional concurrent speakers, recent studies incorporate both overlapped and non-overlapped segments in the training data. However, such…

Cited by 0SourceScholar
2020

Densely Connected Neural Network with Dilated Convolutions for Real-Time Speech Enhancement in The Time Domain

ICASSP 2020accepted

In this work, we propose a fully convolutional neural network for real-time speech enhancement in the time domain. The proposed network is an encoder-decoder based architecture with skip connections. The layers in the encoder and the decoder are followed by densely connected blocks comprising of dil…

Cited by 0SourceScholar
2020

Improving Robustness of Deep Learning Based Monaural Speech Enhancement Against Processing Artifacts

ICASSP 2020accepted

In voice telecommunication, the intelligibility and quality of speech signals can be severely degraded by background noise if the speaker at the transmitting end talks in a noisy environment. Therefore, a speech enhancement system is typically integrated into the transmitter device or the receiver d…

Cited by 12SourceScholar
2019

Complex Spectral Mapping with a Convolutional Recurrent Network for Monaural Speech Enhancement

ICASSP 2019accepted

Phase is important for perceptual quality in speech enhancement. However, it seems intractable to directly estimate phase spectrogram through supervised learning due to lack of clear structure in phase spectrogram. Complex spectral mapping aims to estimate the real and imaginary spectrograms of clea…

Cited by 0SourceScholar
2019

Deep Learning Based Phase Reconstruction for Speaker Separation: A Trigonometric Perspective

ICASSP 2019accepted

This study investigates phase reconstruction for deep learning based monaural talker-independent speaker separation in the short-time Fourier transform (STFT) domain. The key observation is that, for a mixture oftwo sources, with their magnitudes accurately estimated and under a geometric constraint…

Cited by 0SourceScholar
2019

Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios

ICASSP 2019accepted

In mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background n…

Cited by 0SourceScholar
2019

Robust Sparse Multichannel Active Noise Control

ICASSP 2019accepted

Multichannel active noise control (MC-ANC) aims to cancel low-frequency noise in an enclosure. If noise sources are distributed sparsely in space, adding an ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</inf> -norm constraint to the standard MC-AN…

Cited by 0SourceScholar
2019

TCNN: Temporal Convolutional Neural Network for Real-time Speech Enhancement in the Time Domain

ICASSP 2019accepted

This work proposes a fully convolutional neural network (CNN) for real-time speech enhancement in the time domain. The proposed CNN is an encoder-decoder based architecture with an additional temporal convolutional module (TCM) inserted between the encoder and the decoder. We call this architecture…

Cited by 0SourceScholar
2018

Gated Residual Networks with Dilated Convolutions for Supervised Speech Separation

ICASSP 2018accepted

In supervised speech separation, deep neural networks (DNNs) are typically employed to predict an ideal time-frequency (T-F) mask in order to remove background interference. However, the performance of DNNs is frequently degraded for untrained noises and speakers. Inspired by recent research on dila…

Cited by 0SourceScholar
2018

Late Reverberation Suppression Using Recurrent Neural Networks with Long Short-Term Memory

ICASSP 2018accepted

Human speech is usually distorted by room reverberation. These corruptions degrade speech quality and intelligibility, especially under a long reverberation time, and they also pose a serious problem for many speech-related applications such as automatic speech recognition. In this paper, we propose…

Cited by 0SourceScholar
2018

Mask Weighted Stft Ratios for Relative Transfer Function Estimation and ITS Application to Robust ASR

ICASSP 2018accepted

Deep learning based single-channel time-frequency (T-F) masking has shown considerable potential for beamforming and robust ASR. This paper proposes a simple but novel relative transfer function (RTF) estimation algorithm for microphone arrays, where the RTF between a reference signal and a non-refe…

Cited by 0SourceScholar
2018

On Spatial Features for Supervised Speech Separation and its Application to Beamforming and Robust ASR

ICASSP 2018accepted

This study integrates complementary spectral and spatial information to elevate deep learning based time-frequency masking and acoustic beamforming. Coherence and directional features are designed as additional input features for deep neural network training to remove diffuse noise and other directi…

Cited by 0SourceScholar
2018

Utterance-Wise Recurrent Dropout and Iterative Speaker Adaptation for Robust Monaural Speech Recognition

ICASSP 2018accepted

This study addresses monaural (single-microphone) automatic speech recognition (ASR) in adverse acoustic conditions. Our study builds on a state-of-the-art monaural robust ASR method that uses a wide residual network with bidirectional long short-term memory (BLSTM). We propose a novel utterance-wis…

Cited by 12SourceScholar
2017

A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASR

ICASSP 2017accepted

We propose a speech enhancement algorithm based on single- and multi-microphone processing techniques. The core of the algorithm estimates a time-frequency mask which represents the target speech and use masking-based beamforming to enhance corrupted speech. Specifically, in single-microphone proces…

Cited by 0SourceScholar