← Search

Chenglin Xu

10 accepted papers

2026

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

ICML 2026poster

Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document level and focus mainly on high-resource languages, leaving many fine-grained challenges insufficiently evaluated. To address this gap, we present MGAL, the firs…

Cited by 0SourceScholar
2022

L-SpEx: Localized Target Speaker Extraction

ICASSP 2022accepted

Speaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s…

Cited by 0SourceScholar
2022

Multi-Stage and Multi-Loss Training for Fullband Non-Personalized and Personalized Speech Enhancement

ICASSP 2022accepted

Deep learning-based wideband (16kHz) speech enhancement approaches have surpassed traditional methods. This work further extends the existing wideband systems to enable full-band (48kHz) speech enhancement while simultaneously ensuring automatic speech recognition compatibility and optionally, perso…

Cited by 0SourceScholar
2021

Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial Training

ICASSP 2021accepted

Neural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the…

Cited by 0SourceScholar
2021

Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals

ICASSP 2021accepted

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted s…

Cited by 0SourceScholar
2021

Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion Recognition

ICASSP 2021accepted

Convolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spe…

Cited by 0SourceScholar
2020

Time-Domain Neural Network Approach for Speech Bandwidth Extension

ICASSP 2020accepted

In this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of informat…

Cited by 0SourceScholar
2019

Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss

ICASSP 2019accepted

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and cons…

Cited by 0SourceScholar
2018

Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTM

ICASSP 2018accepted

Utterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. Howev…

Cited by 0SourceScholar