← Search

Xugang Lu

17 accepted papers

2025

Continual Unsupervised Domain Adaptation for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models…

Cited by 0SourceScholar
2024

Evaluation of an Improved Ultrasonic Imaging Helmet for Observing Articulatory Data

ICASSP 2024accepted

Ultrasonic imaging is one of the most popular methods for tracking tongue motion. Imaging plane shift and contact variation are crucial factors affecting the consistency of the obtained ultrasonic images. To solve this issue, researchers proposed many different helmets. In this study, we propose an…

Cited by 0SourceScholar
2024

Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASR

ICASSP 2024accepted

Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowled…

Cited by 0SourceScholar
2024

Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition

ICASSP 2024accepted

In the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domai…

Cited by 0SourceScholar
2023

Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification

ICASSP 2023accepted

Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has diffi…

Cited by 0SourceScholar
2022

CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization

ICASSP 2022accepted

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-s…

Cited by 0SourceScholar
2021

Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification

ICASSP 2021accepted

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch prob…

Cited by 0SourceScholar
2021

Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal Transport

NeurIPS 2021poster

This paper presents a novel discriminator-constrained optimal transport network (DOTN) that performs unsupervised domain adaptation for speech enhancement (SE), which is an essential regression task in speech processing. The DOTN aims to estimate clean references of noisy speech in a target domain,…

2020

Robust Unsupervised Neural Machine Translation with Adversarial Denoising Training

COLING 2020main

Unsupervised neural machine translation (UNMT) has recently attracted great interest in the machine translation community. The main advantage of the UNMT lies in its easy collection of required large training text sentences while with only a slightly worse performance than supervised neural machine…

2020

Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement

ICASSP 2020accepted

Nonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for…

Cited by 0SourceScholar
2019

Interactive Learning of Teacher-student Model for Short Utterance Spoken Language Identification

ICASSP 2019accepted

Short utterance-based spoken language identification (LID) is a challenging task due to the large variation of its feature representation. Improving feature representation of short utterances using a teacher-student method has been shown its effectiveness for LID tasks. However, conventional teacher…

Cited by 0SourceScholar
2018

Speech Dereverberation Based on Integrated Deep and Ensemble Learning Algorithm

ICASSP 2018accepted

Reverberation, which is generally caused by sound reflections from walls, ceilings, and floors, can result in severe performance degradation of acoustic applications. Due to a complicated combination of attenuation and time-delay effects, the reverberation property is difficult to characterize, and…

Cited by 0SourceScholar
2017

Minimum Bayes risk training of CTC acoustic models in maximum a posteriori based decoding framework

ICASSP 2017accepted

When using connectionist temporal classification (CTC) based acoustic models (AMs) for large vocabulary continuous speech recognition (LVCSR), most previous studies have used a naive interpolation of the CTC-AM score and an additional language model score, although there is no theoretical justificat…

Cited by 0SourceScholar
2017

Semi-supervised ensemble DNN acoustic model training

ICASSP 2017accepted

It is very important to exploit abundant unlabeled speech for improving the acoustic model training in automatic speech recognition (ASR). Semi-supervised training methods incorporate unlabeled data in addition to labeled data to enhance the model training, but it encounters the error-prone label pr…

Cited by 0SourceScholar
2016

Bottleneck linear transformation network adaptation for speaker adaptive training-based hybrid DNN-HMM speech recognizer

ICASSP 2016accepted

Recently, a Hybrid DNN-HMM recognizer trained with the Speaker Adaptive Training (SAT) concept was successfully modified to a more effective speaker-adaptation-oriented recognizer whose DNN front-end adopted a Linear Transformation Network (LTN) Speaker Dependent (SD) module. However, the size of SD…

Cited by 0SourceScholar
2016

Local fisher discriminant analysis for spoken language identification

ICASSP 2016accepted

I-vector is a state-of-the-art technique widely used in spoken language identification systems. Since i-vectors include total variability factors, discriminant analysis methods have been introduced to find the most discriminative features while removing the undesired variables for language identific…

Cited by 0SourceScholar
2015

Speaker adaptive training for deep neural networks embedding linear transformation networks

ICASSP 2015accepted

Recently, a novel speaker adaptation method was proposed that applied the Speaker Adaptive Training (SAT) concept to a speech recognizer consisting of a Deep Neural Network (DNN) and a Hidden Markov Model (HMM), and its utility was demonstrated. This method implements the SAT scheme by allocating on…

Cited by 0SourceScholar