← Search

Zexu Pan

19 accepted papers

2026

BEYOND LIPS: INTEGRATING GESTURE AND LIP CUES FOR ROBUST AUDIO-VISUAL SPEAKER EXTRACTION

ICASSP 2026poster

Most audio-visual speaker extraction methods rely on synchronized lip recording to isolate the speech of a target speaker from a multi-talker mixture. However, in natural human communication, co-speech gestures are also temporally aligned with speech, often emphasizing specific words or syllables. T…

Cited by 0SourcePDFScholar
2026

PTSE-T: PRESENTATION TARGET SPEAKER EXTRACTION USING UNALIGNED TEXT CUES

ICASSP 2026poster

Target Speaker Extraction (TSE) aims to extract the clean speech of the target speaker in an audio mixture, eliminating irrelevant background noise and speech. While prior work has explored various auxiliary cues including pre-recorded speech, visual information, and spatial information, the acquisi…

Cited by 0SourcePDFScholar
2025

Conditional Latent Diffusion-Based Speech Enhancement via Dual Context Learning

ICASSP 2025accepted

Recently, the application of diffusion probabilistic models has advanced speech enhancement through generative approaches. However, existing diffusion-based methods have focused on the generation process in high-dimensional waveform or spectral domains, leading to increased generation complexity and…

Cited by 0SourceScholar
2025

HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution

ICASSP 2025accepted

The application of generative adversarial networks (GANs) has recently advanced speech super-resolution (SR) based on intermediate representations like mel-spectrograms. However, existing SR methods that typically rely on independently trained and concatenated networks may lead to inconsistent repre…

Cited by 7SourceScholar
2025

Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

ICASSP 2025accepted

The recent rapid development of auditory attention decoding (AAD) offers the possibility of using electroencephalography (EEG) as auxiliary information for target speaker extraction. However, effectively modeling long sequences of speech and resolving the identity of the target speaker from EEG sign…

Cited by 0SourceScholar
2025

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

IJCAI 2025

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which h

2025

SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG

ICASSP 2025accepted

Decoding speech from brain signals is a challenging research problem that holds significant importance for studying speech processing in the brain. Although breakthroughs have been made in reconstructing the mel spectrograms of audio stimuli perceived by subjects at the word or letter level using no…

Cited by 0SourceScholar
2025

Speech Separation for Low-Resource Languages

ICASSP 2025accepted

Speech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can lev…

Cited by 0SourceScholar
2024

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech

ICASSP 2024accepted

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small…

Cited by 0SourceScholar
2024

GLMB 3D Speaker Tracking with Video-Assisted Multi-Channel Audio Optimization Functions

ICASSP 2024accepted

Speaker tracking plays a significant role in numerous real-world human robot interaction (HRI) applications. In recent years, there has been a growing interest in utilizing multi-sensory information, such as complementary audio and visual signals, to address the challenges of speaker tracking. Despi…

Cited by 0SourceScholar
2024

Generation or Replication: Auscultating Audio Latent Diffusion Models

ICASSP 2024accepted

The introduction of audio latent diffusion models possessing the ability to generate realistic sound clips on demand from a text description has the potential to revolutionize how we work with audio. In this work, we make an initial attempt at understanding the inner workings of audio latent diffusi…

Cited by 0SourceScholar
2024

LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing Mechanism

ICASSP 2024accepted

The prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization a…

Cited by 0SourceScholar
2024

NIIRF: Neural IIR Filter Field for HRTF Upsampling and Personalization

ICASSP 2024accepted

Head-related transfer functions (HRTFs) are important for immersive audio, and their spatial interpolation has been studied to upsample finite measurements. Recently, neural fields (NFs) which map from sound source direction to HRTF have gained attention. Existing NF-based methods focused on estimat…

Cited by 0SourceScholar
2024

NeuroHeed+: Improving Neuro-Steered Speaker Extraction with Joint Auditory Attention Detection

ICASSP 2024accepted

Neuro-steered speaker extraction aims to extract the listener’s brainattended speech signal from a multi-talker speech signal, in which the attention is derived from the cortical activity. This activity is usually recorded using electroencephalography (EEG) devices. Though promising, current methods…

Cited by 0SourceScholar
2024

Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition

AAAI 2024technical

Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a vid…

Cited by 11SourcePDFScholar
2023

ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding Inpainting

ICASSP 2023accepted

The speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video cl…

Cited by 0SourceScholar
2021

Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism

ICASSP 2021accepted

Most of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker local…

Cited by 0SourceScholar