← Search

Shilei Zhang

9 accepted papers

2026

OneVoice: One Model, Triple Scenarios—Towards Unified Zero-shot Voice Conversion

IJCAI 2026

Recent progress of voice conversion (VC) has achieved a new milestone in speaker cloning and linguistic preservation. But the field remains fragmented, relying on specialized models for linguistic-preserving, expressive, and singing scenarios. We propose OneVoice, a unified zero-shot framework capab

Cited by 0Scholar
2025

Codec-ASV: Exploring Neural Audio Codec For Speaker Representation Learning

ICASSP 2025accepted

Discrete speech representations have gained significant success in a variety of speech-related tasks. Among these, Neural Audio Codec (NAC), which serves as a compressed form of audio signals, have proven effective in speech AIGC applications. Moreover, we believe that the speaker information can be…

Cited by 0SourceScholar
2025

DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles

COLING 2025main

Human speech exhibits rich and flexible prosodic variations. To address the one-to-many mapping problem from text to prosody in a reasonable and flexible manner, we propose DiffStyleTTS, a multi-speaker acoustic model based on a conditional diffusion module and an improved classifier-free guidance,…

Cited by 2SourcePDFScholar
2025

Efficient Extreme Large-Scale Speaker Verification: Dynamic Active Sub Fully-Connected Layers for Faster Training and Memory Optimization

ICASSP 2025accepted

Using larger scale datasets in the training stage of speaker verification model usually leads to better performance. However, when the speaker number of the training dataset becomes extreme large (e.g., more than 1 million), the training speed and GPU memory demand will become bottlenecks which are…

Cited by 0SourceScholar
2025

Energy-based Model Guided Self-Supervised Learning for Speaker Verification

ICASSP 2025accepted

Self-supervised learning (SSL) has significantly advanced speaker verification, especially in scenarios with limited labeled data. This paper introduces Energy-based Confidence-Aware Distillation (EBCA-DINO), an SSL enhancement for speaker verification that integrates Energy-Based Models (EBMs) into…

Cited by 0SourceScholar
2023

Semi-Supervised Speech Enhancement Based On Speech Purity

ICASSP 2023accepted

We tend to assume most available speech corpora we use are either completely clean or completely noised. However, the reality is most of them are a mix of both. In this paper, we propose a semi-supervised speech enhancement framework to enhance such typical speech datasets. This framework includes a…

Cited by 0SourceScholar
2023

VE-KWS: Visual Modality Enhanced End-to-End Keyword Spotting

ICASSP 2023accepted

The performance of the keyword spotting (KWS) system based on audio modality, commonly measured in false alarms and false rejects, degrades significantly under the far field and noisy conditions. Therefore, audio-visual keyword spotting, which leverages complementary relationships over multiple moda…

Cited by 0SourceScholar
2022

HGCN: Harmonic Gated Compensation Network for Speech Enhancement

ICASSP 2022accepted

Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially masked by noise. To tackle this challenge, we propose a harmon…

Cited by 0SourceScholar
2022

Harmonic Gated Compensation Network Plus for ICASSP 2022 DNS Challenge

ICASSP 2022accepted

The harmonic structure of speech is resistant to noise, but the harmonics may still be partially masked by noise. Therefore, we previously proposed a harmonic gated compensation network (HGCN) to predict the full harmonic locations based on the unmasked harmonics and process the result of a coarse e…

Cited by 0SourceScholar