← Search

Xueliang Zhang

33 accepted papers

2026

Remote Sensing Image Super-Resolution for Imbalanced Textures: A Texture-Aware Diffusion Framework

CVPR 2026

Generative diffusion priors have recently achieved state-of-the-art performance in natural image super-resolution, demonstrating a powerful capability to synthesize photorealistic details. However, their direct application to remote sensing image super-resolution (RSISR) reveals significant shortcom

Cited by 0SourcecodeScholar
2026

Transforming Weather Data from Pixel to Latent Space

ICML 2026oral

The increasing impact of climate change and extreme weather events has spurred growing interest in deep learning for weather research. However, existing studies often rely on weather data in pixel space, which presents several challenges such as smooth outputs in model outputs, limited applicability…

Cited by 0SourceScholar
2025

Enhancing Multi-Channel Speech with Limited Microphones via Spherical Harmonic Transform

ICASSP 2025accepted

The performance of traditional beamforming algorithms is influenced by the number of microphones, with performance improving as the number increases. However, in practice, the number of microphones is often limited. In this paper, we propose a novel virtual microphone estimation method that combines…

Cited by 0SourceScholar
2025

Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models

NeurIPS 2025poster

Vision-Language Models (VLMs) have demonstrated great potential in interpreting remote sensing (RS) images through language-guided semantic. However, the effectiveness of these VLMs critically depends on high-quality image-text training data that captures rich semantic relationships between visual c…

Cited by 0SourceScholar
2025

Vector Quantized Diffusion Model Based Speech Bandwidth Extension

ICASSP 2025accepted

Recent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that util…

Cited by 0SourceScholar
2024

3S-TSE: Efficient Three-Stage Target Speaker Extraction for Real-Time and Low-Resource Applications

ICASSP 2024accepted

Target speaker extraction (TSE) aims to isolate a specific voice from multiple mixed speakers relying on a registerd sample. Since voiceprint features usually vary greatly, current end-to-end neural networks require large model parameters which are computational intensive and impractical for real-ti…

Cited by 0SourceScholar
2024

Efficient Multi-Channel Speech Enhancement with Spherical Harmonics Injection for Directional Encoding

ICASSP 2024accepted

Multi-channel speech enhancement extracts speech using multiple microphones that capture spatial cues. Effectively utilizing directional information is therefore key. Deep learning shows great potential on multi-channel speech enhancement and often takes short-time Fourier Transform (STFT) as inputs…

Cited by 0SourceScholar
2024

Hierarchical Speaker Representation for Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the…

Cited by 0SourceScholar
2024

Innovative Directional Encoding in Speech Processing: Leveraging Spherical Harmonics Injection for Multi-Channel Speech Enhancement

IJCAI 2024poster

Multi-channel speech enhancement leverages multiple microphones to extract target speech signals amid background noise. Effectively utilizing directional cues is key for robust enhancement. While deep learning shows promise for multi-channel speech processing, most methods operate on short-time Four…

2024

SICRN: Advancing Speech Enhancement through State Space Model and Inplace Convolution Techniques

ICASSP 2024accepted

Speech enhancement aims to improve speech quality and intelligibility, especially in noisy environments where background noise degrades speech signals. Currently, deep learning methods achieve great success in speech enhancement, e.g. the representative convolutional recurrent neural network (CRN) a…

Cited by 0SourceScholar
2023

ICCRN: Inplace Cepstral Convolutional Recurrent Neural Network for Monaural Speech Enhancement

ICASSP 2023accepted

According to the mechanism of speech production, speech can be decomposed into excitation and vocal tract which are sparsely represented in cepstral domain. In this study, we propose a neural network for monaural speech enhancement on time-frequency cepstral space that is implemented by inserting a…

Cited by 0SourceScholar
2023

Speech Enhancement with Intelligent Neural Homomorphic Synthesis

ICASSP 2023accepted

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral ana…

Cited by 0SourceScholar
2022

A Complex Spectral Mapping with Inplace Convolution Recurrent Neural Networks For Acoustic Echo Cancellation

ICASSP 2022accepted

Recently, deep learning is introduced in acoustic echo cancellation (AEC) and achieves remarkable performance. For deep learning-based AEC, the most important problem is generalization ability in diversity scenarios. Different from most methods which process the entire frequency band, we propose inp…

Cited by 0SourceScholar
2022

A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature

ICASSP 2022accepted

There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose…

Cited by 0SourceScholar
2022

Alleviating the Loss-Metric Mismatch in Supervised Single-Channel Speech Enhancement

ICASSP 2022accepted

In this paper, we study the loss-metric mismatch problem of supervised single-channel speech enhancement system. Most of the existing speech enhancement systems achieve unsatisfying performance since their empirically selected loss functions have semantic gaps with the non-differentiable evaluation…

Cited by 0SourceScholar
2022

Attention-Based Fusion for Bone-Conducted and Air-Conducted Speech Enhancement in the Complex Domain

ICASSP 2022accepted

Bone-conduction (BC) microphones capture speech signals by converting the vibrations of the human skull into electrical signals. BC sensors are insensitive to acoustic noise, but limited in bandwidth. On the other hand, conventional or air-conduction (AC) microphones are capable of capturing full-ba…

Cited by 0SourceScholar
2022

DRC-NET: Densely Connected Recurrent Convolutional Neural Network for Speech Dereverberation

ICASSP 2022accepted

Under our previous work on frequency bin-wise independent processing, a dramatic reduction of the computational complexity for recurrent neural networks (RNN) is achieved. So that a massive deployment of RNN in time dimension is realized in this paper, by using the channel-wise long short-term memor…

Cited by 0SourceScholar
2021

Real-Time Speech Enhancement for Mobile Communication Based on Dual-Channel Complex Spectral Mapping

ICASSP 2021accepted

Speech quality and intelligibility can be severely degraded by back-ground noise in mobile communication. In order to attenuate back-ground noise, speech enhancement systems have been integrated into mobile phones, and a microphone array is typically deployed to improve the enhancement performance.…

Cited by 4SourceScholar
2019

A Robust Text-independent Speaker Verification Method Based on Speech Separation and Deep Speaker

ICASSP 2019accepted

Recently, deep neural networks (DNNs) have achieved incredible performance in speaker verification. However, most of which remains sensitive to environment noise. In this paper, we propose an end-to-end speaker verification framework to enhance the robustness against background noise. The proposed f…

Cited by 0SourceScholar
2019

Real-time Speech Enhancement Using an Efficient Convolutional Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios

ICASSP 2019accepted

In mobile speech communication, the quality and intelligibility of the received speech can be severely degraded by background noise if the far-end talker is in an adverse acoustic environment. Therefore, speech enhancement algorithms are typically integrated into mobile phones to remove background n…

Cited by 0SourceScholar
2018

Training Supervised Speech Separation System to Improve STOI and PESQ Directly

ICASSP 2018accepted

Supervised speech separation methods train learning machine to cast the noisy speech to the target clean speech. Most of them use mean-square error (MSE) as loss function. However, MSE is not the perfect choice because it doesn't match the human auditory perception. Short-time objective intelligibil…

Cited by 0SourceScholar
2017

A speech enhancement algorithm by iterating single- and multi-microphone processing and its application to robust ASR

ICASSP 2017accepted

We propose a speech enhancement algorithm based on single- and multi-microphone processing techniques. The core of the algorithm estimates a time-frequency mask which represents the target speech and use masking-based beamforming to enhance corrupted speech. Specifically, in single-microphone proces…

Cited by 0SourceScholar
2016

Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation

ICASSP 2016accepted

The targets of speech separation, whether ideal masks or magnitude spectrograms of interest, have prominent spectro-temporal structures. These characteristics are very worthy to be exploited for speech separation, however, they are usually ignored in previous works. In this paper, we use nonnegative…

Cited by 0SourceScholar
2015

A pairwise algorithm for pitch estimation and speech separation using deep stacking network

ICASSP 2015accepted

Pitch information is an important cue for speech separation. However, pitch estimation in noisy condition is also a task as challenging as speech separation. In this paper, we propose a supervised learning architecture which combines these two problems concisely. The proposed algorithm is based on d…

Cited by 0SourceScholar