← Search

Xin Fang

8 accepted papers

2026

Improving Anomalous Sound Detection with Attribute-aware Representation from Domain-adaptive Pre-training

ICASSP 2026poster

Anomalous Sound Detection (ASD) is often formulated as a machine attribute classification task, a strategy necessitated by the common scenario where only normal data is available for training. However, the exhaustive collection of machine attribute labels is laborious and impractical. To address the…

Cited by 0SourcePDFScholar
2025

MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation

ICASSP 2025accepted

Sound event localization and detection with source distance estimation (3D SELD) involves not only identifying the sound category and its direction-of-arrival (DOA) but also predicting the source's distance, aiming to provide full information about the sound position. This paper proposes a multi-sta…

Cited by 0SourceScholar
2024

Implicit Enhancement of Target Speaker in Speaker-Adaptive ASR through Efficient Joint Optimization

ICASSP 2024accepted

In multi-speaker scenarios, automatic speech recognition (ASR) models rely on pre-processed audio after speaker separation. However, when the target speaker is not accurately separated, ASR models face limitations in reaching their peak performance. To address this issue, we propose a speaker-adapti…

Cited by 0SourceScholar
2023

A Multi-Scale Feature Aggregation Based Lightweight Network for Audio-Visual Speech Enhancement

ICASSP 2023accepted

Audio-visual speech enhancement (AVSE) was shown to be superior over conventional audio-only counterpart for improving the speech quality. However, most existing AVSE models are heavyweight in the sense of parameter count, which is inappropriate for the deployment and practical applications. In this…

Cited by 0SourceScholar
2023

AST-SED: An Effective Sound Event Detection Method Based on Audio Spectrogram Transformer

ICASSP 2023accepted

In this paper, we propose an effective sound event detection (SED) method based on the audio spectrogram transformer (AST) model, pretrained on the large-scale AudioSet for audio tagging (AT) task, termed AST-SED. Pretrained AST models have recently shown promise on DCASE2022 challenge task4 where t…

Cited by 0SourceScholar
2022

A Noise-Robust Self-Supervised Pre-Training Model Based Speech Representation Learning for Automatic Speech Recognition

ICASSP 2022accepted

Wav2vec2.0 is a popular self-supervised pre-training framework for learning speech representations in the context of automatic speech recognition (ASR). It was shown that wav2vec2.0 has a good robustness against the domain shift, while the noise robustness is still unclear. In this work, we therefor…

Cited by 0SourceScholar
2020

Progressive Multi-Target Network Based Speech Enhancement with Snr-Preselection for Robust Speaker Diarization

ICASSP 2020accepted

In this paper, we design a novel front-end processing system for speaker diarization under realistic conditions with challenging background noises. To cope with diversified environments, we first extend our perviously proposed progressive learning based speech enhancement model by adding multi-task…

Cited by 0SourceScholar
2019

Channel Adversarial Training for Cross-channel Text-independent Speaker Recognition

ICASSP 2019accepted

The conventional speaker recognition frameworks (e.g., the i-vector and CNN-based approach) have been successfully applied to various tasks when the channel of the enrolment dataset is similar to that of the test dataset. However, in real-world applications, mismatch always exists between these two…

Cited by 0SourceScholar