← Search

Xinmeng Xu

13 accepted papers

2026

Neural-Inspired Modeling of Auditory Selection and Compensation for Audio-Visual Speech Separation

ICML 2026poster

Current audio-visual speech separation (AVSS) models typically rely on implicit multimodal fusion, but the absence of explicit modality alignment and reliability modeling often causes semantic misalignment and contaminates speech representations. The brain addresses this with a hierarchy: top-down a…

Cited by 0SourceScholar
2025

Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs

ICASSP 2025accepted

Existing end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality durin…

Cited by 0SourceScholar
2025

FIRING-Net: A filtered feature recycling network for speech enhancement

ICLR 2025poster

Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination…

Cited by 0SourcePDFScholar
2025

Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model

ICASSP 2025accepted

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additi…

Cited by 0SourceScholar
2024

An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning

ICASSP 2024accepted

Speech enhancement is a vital and highly ill-posed problem for many speech downstream tasks. While currently existing deep learning based speech enhancement methods have held state-of-the-art results, they still possess apparent shortcomings in that most of the deep learning based models lack interp…

Cited by 0SourceScholar
2024

Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations

ICASSP 2024accepted

Existing deep learning-based speech enhancement methods only adopt clean speech as positive samples to guide the training of speech enhancement networks while negative samples, i.e., noisy speech, are unexploited. In this paper, we adopt contrastive regularization (CR) built upon contrastive learnin…

Cited by 0SourceScholar
2024

Improving Acoustic Echo Cancellation by Exploring Speech and Echo Affinity with Multi-Head Attention

ICASSP 2024accepted

Deep learning-based approaches formulate acoustic echo cancellation (AEC) as a supervised speech separation task, where the mixture signal and the far-end signal are combined directly before or after the encoding stage. However, the mixture signal and the far-end signal are not integrated sufficient…

Cited by 0SourceScholar
2024

SuperCodec: A Neural Speech Codec with Selective Back-Projection Network

ICASSP 2024accepted

Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction…

Cited by 0SourceScholar
2023

Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer

ICASSP 2023accepted

We propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single…

Cited by 0SourceScholar
2023

Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech Enhancement

AAAI 2023technical

Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, a natural speech contains many fast-changing and relatively briefly acoustic events, therefore, capturing the most informative speech features…

Cited by 11SourcePDFScholar
2022

Improving Dual-Microphone Speech Enhancement by Learning Cross-Channel Features with Multi-Head Attention

ICASSP 2022accepted

Hand-crafted spatial features, such as inter-channel intensity difference (IID) and inter-channel phase difference (IPD), play a fundamental role in recent deep learning based dual-microphone speech enhancement (DMSE) systems. However, learning the mutual relationship between artificially designed s…

Cited by 0SourceScholar
2022

VSEGAN: Visual Speech Enhancement Generative Adversarial Network

ICASSP 2022accepted

Speech enhancement is an essential task of improving speech quality in noise scenario. Several state-of-the-art approaches have introduced visual information for speech enhancement, since the visual aspect of speech is essentially unaffected by acoustic environment. This paper proposes a novel frame…

Cited by 0SourceScholar