← Search

Wenhuan Lu

12 accepted papers

2026

DuoKD: Dual Knowledge Distillation from Large Language Models for Robust Graph Neural Networks

AAAI 2026technical

Graph neural networks (GNNs) have become a dominant modeling paradigm for graph-structured data, and the emergence of large language models (LLMs) has spurred growing interest in integrating external semantic knowledge into GNNs. Current LLM-based GNNs are devoted to extracting semantically similar

Cited by 0SourcePDFScholar
2026

EA-VAE: Learning to Reconstruct Dysarthric Speech via Variational Autoencoder with Encoding Alignment

AAAI 2026technical

Dysarthric speech reconstruction (DSR) aims to enhance the intelligibility of dysarthric speech. Compared with normal speech, the dysarthric speech is characterized by its pathological features, including discontinuous pronunciation, slow speech, hoarseness, and improper pauses. Significant disparit

Cited by 0SourcePDFScholar
2025

Adaptive Multi-Scale Local Correction for Semi-Supervised 3D Medical Image Segmentation

ICASSP 2025accepted

In recent years, semi-supervised 3D medical image segmentation has gained significant attention. However, current methods often struggle with multi-scale voxel differences and overlook the importance of loss weight balancing. To address these issues, we propose an adaptive multi-scale local correcti…

Cited by 0SourceScholar
2025

Continual Unsupervised Domain Adaptation for Audio Deepfake Detection

ICASSP 2025accepted

Audio deepfake detection (ADD) aims to verify the authenticity of audio. However, its performance declines sharply when facing significant domain discrepancies caused by unknown datasets. Unsupervised domain adaptation (UDA) has been applied to mitigate domain mismatch. However, as generative models…

Cited by 0SourceScholar
2024

EEG-Based Fast Auditory Attention Detection in Real-Life Scenarios Using Time-Frequency Attention Mechanism

ICASSP 2024accepted

Auditory attention detection (AAD) based on electroencephalogram (EEG) helps recognize the target speaker in a cocktail party scenario, advancing auditory brain-computer interface development. Previous EEG studies on AAD were largely based on data collected in laboratory settings. In this study, we…

Cited by 0SourceScholar
2024

Self-Supervised Domain Exploration with an Optimal Transport Regularization for Open Set Cross-Domain Speech Emotion Recognition

ICASSP 2024accepted

In the tasks of domain adaptation (DA) for speech emotion recognition (SER), self-supervised learning (SSL) algorithms could effectively explore domain and structural information from target domain samples, thereby mitigating domain discrepancies. However, in a general setting, when the target domai…

Cited by 0SourceScholar
2023

Optimal Transport with a Diversified Memory Bank for Cross-Domain Speaker Verification

ICASSP 2023accepted

Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers' probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has diffi…

Cited by 0SourceScholar
2022

CS-REP: Making Speaker Verification Networks Embracing Re-Parameterization

ICASSP 2022accepted

Automatic speaker verification (ASV) systems, which determine whether two speeches are from the same speaker, mainly focus on verification accuracy while ignoring inference speed. However, in real applications, both inference speed and verification accuracy are essential. This study proposes cross-s…

Cited by 0SourceScholar
2022

Joint and Adversarial Training with ASR for Expressive Speech Synthesis

ICASSP 2022accepted

Style modeling is an important issue and has been proposed in expressive speech synthesis. In existing unsupervised methods, the style encoder extracts the latent representation from the reference audio as style information. However, the style information extracted from the style encoder will entang…

Cited by 0SourceScholar
2021

Zero-Shot Voice Conversion with Adjusted Speaker Embeddings and Simple Acoustic Features

ICASSP 2021accepted

Zero-shot voice conversion (VC) where both source and target speakers are unseen in the training dataset has become a new research direction. Using speaker embeddings instead of one-hot vectors to represent speaker identity is a key point, which makes VC models work on unseen speakers. In our work,…

Cited by 0SourceScholar
2019

Breast Cancer Detection Based on Merging Four Modes MRI Using Convolutional Neural Networks

ICASSP 2019accepted

The objective of the study is to develop a framework for automatic breast cancer detection with merging four imaging modes. Attempts were made for tumor classification and segmentation; using a multi-parametric Magnetic Resonance Imaging (MRI) method on breast tumors. MRI data of the breast were obt…

Cited by 0SourceScholar