← Search

Shengkui Zhao

18 accepted papers

2026

BEYOND LIPS: INTEGRATING GESTURE AND LIP CUES FOR ROBUST AUDIO-VISUAL SPEAKER EXTRACTION

ICASSP 2026poster

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. T…

Cited by 0SourcePDFScholar
2025

Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning

ICASSP 2025accepted

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and…

Cited by 0SourceScholar
2025

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

ICASSP 2025accepted

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent repre…

Cited by 7SourceScholar
2024

Are Soft Prompts Good Zero-Shot Learners for Speech Recognition?

ICASSP 2024accepted

Large self-supervised pre-trained speech models require computationally expensive fine-tuning for downstream tasks. Soft prompt tuning offers a simple parameter-efficient alternative by utilizing minimal soft prompt guidance, enhancing portability while also maintaining competitive performance. Howe…

Cited by 0SourceScholar
2024

MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation

ICASSP 2024accepted

Our previously proposed MossFormer has achieved promising performance in monaural speech separation. However, it predominantly adopts a self-attention-based MossFormer module, which tends to emphasize longer-range, coarser-scale dependencies, with a deficiency in effectively modelling finer-scale re…

Cited by 0SourceScholar
2024

SPGM: Prioritizing Local Features for Enhanced Speech Separation Performance

ICASSP 2024accepted

Dual-path is a popular architecture for speech separation models (e.g. Sepformer) which splits long sequences into overlapping chunks for its intra- and inter-blocks that separately model intra-chunk local features and inter-chunk global relationships. However, it has been found that inter-blocks, w…

Cited by 0SourceScholar
2023

D2Former: A Fully Complex Dual-Path Dual-Decoder Conformer Network Using Joint Complex Masking and Complex Spectral Mapping for Monaural Speech Enhancement

ICASSP 2023accepted

Monaural speech enhancement has been widely studied using real networks in the time-frequency (TF) domain. However, the input and the target are naturally complex-valued in the TF domain, a fully complex network is highly desirable for effectively learning the feature representation and modelling th…

Cited by 0SourceScholar
2023

MossFormer: Pushing the Performance Limit of Monaural Speech Separation Using Gated Single-Head Transformer with Convolution-Augmented Joint Self-Attentions

ICASSP 2023accepted

Transformer based models have provided significant performance improvements in monaural speech separation. However, there is still a performance gap compared to a recent proposed upper bound. The major limitation of the current dual-path Transformer models is the inefficient modelling of long-range…

Cited by 0SourceScholar
2022

End-to-End Complex-Valued Multidilated Convolutional Neural Network for Joint Acoustic Echo Cancellation and Noise Suppression

ICASSP 2022accepted

Echo and noise suppression is an integral part of a full-duplex communication system. Many recent acoustic echo cancellation (AEC) systems rely on a separate adaptive filtering module for linear echo suppression and a neural module for residual echo suppression. However, in practice, adaptive filter…

Cited by 0SourceScholar
2022

FRCRN: Boosting Feature Representation Using Frequency Recurrence for Monaural Speech Enhancement

ICASSP 2022accepted

Convolutional recurrent networks (CRN) integrating a convolutional encoder-decoder (CED) structure and a recurrent structure have achieved promising performance for monaural speech enhancement. However, feature representation across frequency context is highly constrained due to limited receptive fi…

Cited by 0SourceScholar
2021

Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses

ICASSP 2021accepted

Deep complex U-Net structure and convolutional recurrent network (CRN) structure achieve state-of-the-art performance for monaural speech enhancement. Both deep complex U-Net and CRN are encoder and decoder structures with skip connections, which heavily rely on the representation power of the compl…

Cited by 0SourceScholar
2021

Towards Natural and Controllable Cross-Lingual Voice Conversion Based on Neural TTS Model and Phonetic Posteriorgram

ICASSP 2021accepted

Cross-lingual voice conversion (VC) is an important and challenging problem due to significant mismatches of the phonetic set and the speech prosody of different languages. In this paper, we build upon the neural text-to-speech (TTS) model, i.e., FastSpeech, and LPCNet neural vocoder to design a new…

Cited by 0SourceScholar
2017

A novel sparse model for multi-source localization using distributed microphone array

ICASSP 2017accepted

When distances between microphone pairs are larger than the half-wavelength of signals, source localization methods using cross-correlation such as time-difference-of-arrival (TDOA), steered response power (SRP) are commonly used in practice. We present here a novel model that expresses microphone p…

Cited by 3SourceScholar
2017

On time-frequency mask estimation for MVDR beamforming with application in robust speech recognition

ICASSP 2017accepted

Acoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estim…

Cited by 0SourceScholar
2016

An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sources

ICASSP 2016accepted

This paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector c…

Cited by 0SourceScholar
2016

Large region acoustic source mapping: A generalized sparse constrained deconvolution approach

ICASSP 2016accepted

This paper presents a generalized multiple-point sparse constrained deconvolution approach for mapping acoustic noise sources in large regions using a movable array. Extended from our previous MPSC-DAMAS approach, we first derive a generalized inverse problem relating to the source powers and the ar…

Cited by 0SourceScholar
2015

A learning-based approach to direction of arrival estimation in noisy and reverberant environments

ICASSP 2015accepted

This paper presents a learning-based approach to the task of direction of arrival estimation (DOA) from microphone array input. Traditional signal processing methods such as the classic least square (LS) method rely on strong assumptions on signal models and accurate estimations of time delay of arr…

Cited by 0SourceScholar