← Search

Zhaoxi Mu

5 accepted papers

2024

Self-Supervised Disentangled Representation Learning for Robust Target Speech Extraction

AAAI 2024technical

Speech signals are inherently complex as they encompass both global acoustic characteristics and local semantic information. However, in the task of target speech extraction, certain elements of global and local semantic information in the reference speech, which are irrelevant to speaker identity,…

Cited by 8SourcePDFScholar
2024

Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction

IJCAI 2024poster

The integration of visual cues has revitalized the performance of the target speech extraction task, elevating it to the forefront of the field. Nevertheless, this multi-modal learning paradigm often encounters the challenge of modality imbalance. In audio-visual target speech extraction tasks, the…

Cited by 3SourcePDFScholar
2023

A Multi-Stage Triple-Path Method For Speech Separation in Noisy and Reverberant Environments

ICASSP 2023accepted

In noisy and reverberant environments, the performance of deep learning-based speech separation methods drops dramatically because previous methods are not designed and optimized for such situations. To address this issue, we propose a multi-stage end-to-end learning method that decouples the diffic…

Cited by 0SourceScholar
2023

Multi-Dimensional and Multi-Scale Modeling for Speech Separation Optimized by Discriminative Learning

ICASSP 2023accepted

Transformer has shown advanced performance in speech separation, benefiting from its ability to capture global features. However, capturing local features and channel information of audio sequences in speech separation is equally important. In this paper, we present a novel approach named Intra-SE-C…

Cited by 0SourceScholar